A data processing method and apparatus

By calculating the numerical information and similarity prediction of the entities to be processed, the problems of low data audit efficiency and poor user experience in e-commerce are solved, and efficient data audit and user experience improvement are achieved.

CN112101399BActive Publication Date: 2025-05-27BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910528709.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-18
Publication Date
2025-05-27
Estimated Expiration
2039-06-18

AI Technical Summary

Technical Problem

In e-commerce, the data review is inefficient, large workload, long time, and relies on manual experience, which makes it easy to have audit errors, resulting in poor user experience.

Method used

By calculating the numerical information of the entity to be processed, the target display entity with the smallest difference is found in the display entity linked list, and check whether there is duplicate entity data based on the preset value and similarity prediction.

Benefits of technology

It improves data audit efficiency, reduces workload and time, reduces dependence on manual experience, reduces audit error rate, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112101399B_ABST
    Figure CN112101399B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method and apparatus, relating to the field of computer technology. A specific embodiment of the method includes: calculating numerical information corresponding to a to-be-processed entity according to the input to-be-processed entity data; searching for a target display entity in the display entity linked list with the smallest difference between the corresponding numerical information and the numerical information corresponding to the to-be-processed entity; comparing the difference between the numerical information corresponding to the target display entity and the to-be-processed entity with a preset value. If it is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the to-be-processed entity in the display entity linked list; if it is less than the preset value, the similarity between the target display entity and the to-be-processed entity is predicted, and whether there is duplicate entity data of the to-be-processed entity in the display entity linked list is checked according to the similarity. This embodiment can improve the data review efficiency, reduce the review workload, shorten the review time, reduce the dependence on manual experience, reduce review errors, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a data processing method and apparatus. Background Art

[0002] With the development of current e-commerce towards the integration of online and offline, more and more offline entities need to be displayed online. Offline entities often cooperate with multiple pallet providers (or distributors). At the same time, e-commerce companies also cooperate with multiple pallet providers. Thus, when e-commerce companies display these entities online, the following situations may occur: the same entity is entered into the system by different pallet providers, and the information may be slightly different, for example, only the name is different. In this way, duplicate entities will appear when displaying these entities, resulting in a poor user experience.

[0003] Currently, data review relies on manual work, or through online user complaints, and then manual review is carried out to mark duplicate entities and not display them online. This method of completely relying on manual work to eliminate duplicate entities requires manual comparison, which not only has low efficiency and heavy workload, but also completely depends on personal experience. If a new person is replaced to do this work, it will also bring review errors. Moreover, manual review will bring a long review time, resulting in a decline in service quality and causing complaints from service providers.

[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:

[0005] The existing solutions have low data review efficiency, heavy review workload, long review time, high dependence on manual experience, and may also have review errors caused by insufficient personnel experience, resulting in a poor user experience. Summary of the Invention

[0006] In view of this, embodiments of the present invention provide a data processing method and apparatus, which can improve data review efficiency, reduce review workload, shorten review time, reduce dependence on manual experience, reduce review errors, and improve user experience.

[0007] To achieve the above object, according to one aspect of the embodiments of the present invention, a data processing method is provided.

[0008] A data processing method, comprising: calculating numerical information corresponding to a to-be-processed entity according to the input to-be-processed entity data; searching for a target display entity in a display entity linked list with the smallest difference between the corresponding numerical information and the numerical information corresponding to the to-be-processed entity, wherein the display entity linked list stores display entity data, the display entity is an existing entity for display, and each display entity has its own corresponding numerical information; comparing the difference between the numerical information corresponding to the target display entity and the to-be-processed entity with a preset value, if the difference between the numerical information is greater than or equal to the preset value, determining that there is no duplicate entity data of the to-be-processed entity in the display entity linked list; if the difference between the numerical information is less than the preset value, predicting the similarity between the target display entity and the to-be-processed entity, and checking whether there is duplicate entity data of the to-be-processed entity in the display entity linked list according to the similarity.

[0009] Optionally, the display entity linked list is the main chain of a two-dimensional entity linked list, the two-dimensional entity linked list further includes sub-chains, each sub-chain is a duplicate entity linked list of a display entity, which stores the duplicate entity data of the display entity, and the two-dimensional entity linked list is obtained by the following method: selecting specific entities from each existing entity according to a preset rule; calculating the numerical information corresponding to each existing entity, and sorting each existing entity in ascending order of the numerical information to obtain an existing entity linked list; according to the numerical information corresponding to each existing entity, converting the existing entity linked list into a two-dimensional linked list, and the conversion rule is: using the specific entity as the first existing entity of the main chain of the two-dimensional linked list, and successively determining the subsequent existing entities of the main chain of the two-dimensional linked list along the existing entity linked list, and generating sub-chains for each existing entity of the main chain of the two-dimensional linked list, so that the difference between the numerical information corresponding to adjacent existing entities on the main chain of the two-dimensional linked list is greater than or equal to the preset value, and the difference between the numerical information corresponding to the existing entity on the main chain of the two-dimensional linked list and the existing entity of its sub-chain is less than the preset value; generating training data according to the sub-chains of the two-dimensional linked list, training a machine learning algorithm model, and obtaining non-duplicate entities of all sub-chains of the two-dimensional linked list during the training process, the non-duplicate entities of the sub-chain are entities in the sub-chain that are not duplicate with the existing entity on the corresponding main chain of the two-dimensional linked list, and the machine learning algorithm model is a model supporting classification; changing the non-duplicate entities of all sub-chains to the main chain of the two-dimensional linked list, and the obtained new two-dimensional linked list is the two-dimensional entity linked list.

[0010] Optionally, the step of selecting specific entities from each existing entity according to a preset rule includes: obtaining the longitude and latitude information of each existing entity according to the entity address information in the existing entity data; sorting each existing entity according to the longitude and latitude information of each existing entity in a predetermined order, and selecting the first existing entity after sorting as the specific entity.

[0011] Optionally, if there is no duplicate entity data of the entity to be processed in the display entity linked list, the entity data to be processed is stored in the display entity linked list for display; if there is duplicate entity data of the entity to be processed in the display entity linked list, the entity data to be processed is stored in the duplicate entity linked list of the target display entity for non-display.

[0012] Optionally, after receiving a take-off line instruction for an entity in the two-dimensional entity linked list, find the storage location of the entity data to be taken off line; if the entity data to be taken off line is stored in the display entity linked list and there is a duplicate entity linked list of the entity to be taken off line in the two-dimensional entity linked list, delete the stored entity data to be taken off line, and move the entity data with the highest user score in the duplicate entity linked list to the original storage location of the entity data to be taken off line; if the entity data to be taken off line is stored in the display entity linked list and there is no duplicate entity linked list of the entity to be taken off line in the two-dimensional entity linked list, or if the entity data to be taken off line is stored in the duplicate entity linked list of a certain display entity, directly delete the stored entity data to be taken off line.

[0013] Optionally, the steps of generating training data according to the sub-chains of the two-dimensional linked list, training a machine learning algorithm model, and obtaining non-duplicate entities of all sub-chains of the two-dimensional linked list during the training process include: calculating the multi-dimensional attribute similarity between the existing entities in all sub-chains and the existing entities in the main chain of the corresponding two-dimensional linked list according to the entity multi-dimensional attribute similarity calculation rule; selecting a first part of sub-chains from all sub-chains, generating first training data according to the first part of sub-chains, and performing first training on the machine learning algorithm model; selecting a second part of sub-chains from the currently remaining sub-chains, performing similarity prediction on the second part of sub-chains by using the machine learning algorithm model trained in the first training, where the similarity prediction includes predicting the similarity between the existing entities in the sub-chain and the existing entities in the corresponding main chain, and selecting some sub-chains from the second part of sub-chains according to the prediction result to generate second training data, and performing second training on the machine learning algorithm model by using the first and second training data; performing the similarity prediction on the finally remaining sub-chains by using the machine learning algorithm model trained in the second training, and selecting some sub-chains from the finally remaining sub-chains according to the prediction result to generate third training data, and performing third training on the machine learning algorithm model by using the first, second, and third training data; where the first, second, and third training data respectively include the multi-dimensional attribute similarity and annotation values between the existing entities in the corresponding sub-chains and the existing entities in the corresponding main chain, and the annotation values include the annotated non-duplicate entities of the corresponding sub-chains; obtaining the non-duplicate entities of all sub-chains according to the annotated non-duplicate entities of each sub-chain and the non-duplicate entities in each sub-chain identified by each prediction result of the machine learning algorithm model.

[0014] Optionally, the step of predicting the similarity between the target display entity and the entity to be processed includes: calculating the multi-dimensional attribute similarity between the target display entity and the entity to be processed according to the entity multi-dimensional attribute similarity calculation rule, and inputting the multi-dimensional attribute similarity into the machine learning algorithm model trained by the third training to predict the similarity between the target display entity and the entity to be processed.

[0015] Optionally, the multi-dimensional attribute similarity includes distance similarity, phone similarity, name similarity, address similarity, and picture similarity. The entity multi-dimensional attribute similarity calculation rule includes calculating the multi-dimensional attribute similarity between a first entity and a second entity in the following manner: the distance similarity = 1 - m / k, where m is the difference between the numerical information corresponding to the first entity and the second entity, k is the preset value, and m / k represents the ratio of m to k. The numerical information corresponding to the first entity is the distance from the first entity to the specific entity, which is calculated using the longitude and latitude information of the first entity and the specific entity; the numerical information corresponding to the second entity is the distance from the second entity to the specific entity, which is calculated using the longitude and latitude information of the second entity and the specific entity; the phone similarity is obtained by comparing the phone numbers of the first entity and the second entity digit by digit, and counting the number of inconsistent digits between the two phone numbers, and looking up the phone similarity value corresponding to the number of digits; the name similarity and the address similarity are obtained through a text-to-vector model; the picture similarity is obtained through a picture similarity matching algorithm model; where the first entity is an existing entity in the sub-chain of the two-dimensional linked list, the second entity is an existing entity in the main chain of the two-dimensional linked list, or the first entity is the entity to be processed, and the second entity is the target display entity.

[0016] According to another aspect of the embodiments of the present invention, a data processing device is provided.

[0017] A data processing device, comprising: a calculation module for calculating numerical information corresponding to a to-be-processed entity according to the input to-be-processed entity data; a search module for searching in a display entity linked list for a target display entity with the smallest difference between the corresponding numerical information and the numerical information corresponding to the to-be-processed entity, the display entity linked list storing display entity data, the display entity being an existing entity for display, and each display entity having its own corresponding numerical information; a screening and processing module for comparing the difference between the numerical information corresponding to the target display entity and the to-be-processed entity with a preset value. If the difference in numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the to-be-processed entity in the display entity linked list; if the difference in numerical information is less than the preset value, the similarity between the target display entity and the to-be-processed entity is predicted, and whether there is duplicate entity data of the to-be-processed entity in the display entity linked list is screened according to the similarity.

[0018] Optionally, the display entity linked list is the main chain of a two-dimensional entity linked list. The two-dimensional entity linked list further includes sub-chains, and each sub-chain is a duplicate entity linked list of a display entity, which stores the duplicate entity data of the display entity. The device further includes a two-dimensional linked list generation module, and obtains the two-dimensional entity linked list through the following method: selecting specific entities from each existing entity according to a preset rule; calculating the numerical information corresponding to each existing entity, and sorting each existing entity in ascending order of the numerical information to obtain an existing entity linked list; according to the numerical information corresponding to each existing entity, converting the existing entity linked list into a two-dimensional linked list, and the conversion rule is: using the specific entity as the first existing entity of the main chain of the two-dimensional linked list, and sequentially determining the subsequent existing entities of the main chain of the two-dimensional linked list along the existing entity linked list, and generating sub-chains for each existing entity of the main chain of the two-dimensional linked list, so that the difference between the numerical information corresponding to adjacent existing entities on the main chain of the two-dimensional linked list is greater than or equal to the preset value, and the difference between the numerical information corresponding to the existing entity on the main chain of the two-dimensional linked list and the existing entity in its sub-chain is less than the preset value; generating training data according to the sub-chains of the two-dimensional linked list, training a machine learning algorithm model, and obtaining non-duplicate entities of all sub-chains of the two-dimensional linked list during the training process. The non-duplicate entity of the sub-chain is an entity in the sub-chain that is not duplicate with the existing entity on the corresponding main chain of the two-dimensional linked list. The machine learning algorithm model is a model for supporting classification; changing the non-duplicate entities of all sub-chains to the main chain of the two-dimensional linked list, and the obtained new two-dimensional linked list is the two-dimensional entity linked list.

[0019] Optionally, it further includes a display processing module, which is used to: if there is no duplicate entity data of the entity to be processed in the display entity linked list, store the entity data to be processed in the display entity linked list for display; if there is duplicate entity data of the entity to be processed in the display entity linked list, store the entity data to be processed in the duplicate entity linked list of the target display entity for non-display.

[0020] Optionally, it further includes a deactivation processing module, which is used to: after receiving a deactivation instruction for an entity in the two-dimensional entity linked list, find the storage location of the entity data to be deactivated; if the entity data to be deactivated is stored in the display entity linked list and there is a duplicate entity linked list of the entity to be deactivated in the two-dimensional entity linked list, delete the stored entity data to be deactivated and move the entity data with the highest user score in the duplicate entity linked list to the original storage location of the entity data to be deactivated; if the entity data to be deactivated is stored in the display entity linked list and there is no duplicate entity linked list of the entity to be deactivated in the two-dimensional entity linked list, or if the entity data to be deactivated is stored in the duplicate entity linked list of a certain display entity, directly delete the stored entity data to be deactivated.

[0021] Optionally, the troubleshooting processing module includes a similarity prediction sub-module, which is used to: calculate the multi-dimensional attribute similarity between the target display entity and the entity to be processed according to the entity multi-dimensional attribute similarity calculation rule, and input the multi-dimensional attribute similarity into the trained machine learning algorithm model to predict the similarity between the target display entity and the entity to be processed.

[0022] According to another aspect of the embodiments of the present invention, an electronic device is provided.

[0023] An electronic device includes: one or more processors; a memory for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the data processing method provided by the present invention.

[0024] According to another aspect of the embodiments of the present invention, a computer-readable medium is provided.

[0025] A computer-readable medium has a computer program stored thereon, and when the program is executed by a processor, it implements the data processing method provided by the present invention.

[0026] One embodiment of the above invention has the following advantages or beneficial effects: calculating numerical information corresponding to a to-be-processed entity according to the input to-be-processed entity data; searching for a target display entity in the display entity linked list with the smallest difference between the corresponding numerical information and the numerical information of the to-be-processed entity; comparing the difference between the numerical information of the target display entity and the to-be-processed entity with a preset value. If the difference in numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the to-be-processed entity in the display entity linked list; if the difference in numerical information is less than the preset value, predicting the similarity between the target display entity and the to-be-processed entity, and checking whether there is duplicate entity data of the to-be-processed entity in the display entity linked list according to the similarity. This implementation can improve data review efficiency, reduce review workload, shorten review time, reduce reliance on manual experience, reduce review errors, and improve user experience.

[0027] The further effects of the above non-conventional optional methods will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0029] Figure 1 is a schematic diagram of the main steps of the data processing method according to an embodiment of the present invention;

[0030] Figure 2 is a schematic diagram of the data processing flow according to an embodiment of the present invention;

[0031] Figure 3 is a schematic diagram of a store list according to an embodiment of the present invention;

[0032] Figure 4 is a schematic diagram of a two-dimensional linked list according to an embodiment of the present invention;

[0033] Figure 5 is a schematic diagram of the identification of duplicate stores in a sub-chain according to an embodiment of the present invention;

[0034] Figure 6 is a schematic diagram of a two-dimensional store linked list according to an embodiment of the present invention;

[0035] Figure 7 is a schematic diagram of the main modules of the data processing device according to an embodiment of the present invention;

[0036] Figure 8 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0037] Figure 9 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0038] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0039] Those skilled in the art know that the implementation manners of the present invention can be realized as a system, a device, an equipment, a method, or a computer program product. Therefore, the present disclosure can be specifically realized in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0040] Figure 1 is a schematic diagram of the main steps of the data processing method according to an embodiment of the present invention.

[0041] As Figure 1 shown, the data processing method of the embodiment of the present invention mainly includes the following steps S101 to S103.

[0042] Step S101: Calculate numerical information corresponding to the entity to be processed according to the input entity data to be processed.

[0043] The numerical information corresponding to the entity to be processed is, for example, the distance value corresponding to the entity to be processed, and the distance value specifically refers to the distance from the entity to be processed to a specific entity. The specific entity is the first entity in the display entity linked list.

[0044] Step S102: Search for a target display entity in the display entity linked list whose difference from the numerical information corresponding to the entity to be processed is the smallest.

[0045] The display entity linked list stores display entity data. The display entity is an existing entity for display, and each display entity has its own corresponding numerical information.

[0046] The numerical information corresponding to the display entity is, for example, the distance value corresponding to the display entity, and the distance value specifically refers to the distance from the display entity to a specific entity.

[0047] Step S103: Compare the difference between the numerical information corresponding to the target display entity and the entity to be processed with a preset value. If the difference in the numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the entity to be processed in the display entity linked list; if the difference in the numerical information is less than the preset value, predict the similarity between the target display entity and the entity to be processed, and check whether there is duplicate entity data of the entity to be processed in the display entity linked list according to the similarity.

[0048] The entity in the embodiment of the present invention can be a store, an offline network point, a scenic spot, a specific place, etc., which has an actual site location offline and can be promoted and displayed online. Online means on the Internet, and offline, relative to online, refers to what actually exists and is not on the Internet. The data processing method in the embodiment of the present invention can perform a series of processes on the input entity data to be processed, mainly including checking for entities that are duplicates of the entity to be processed from the display entities, and further processing the entity data to be processed according to the check result, specifically including displaying the entity to be processed or not displaying the entity to be processed.

[0049] Taking the entity as a store as an example below, combined with Figure 2 the data processing process, the above steps will be introduced in detail. In the scenario where the entity is a store, as Figure 2 shown, the data processing process in the embodiment of the present invention mainly includes the following steps S201 to S203.

[0050] Step S201: Calculate the distance value corresponding to the store to be processed according to the input store data to be processed.

[0051] The store data may include store address information, store name, store phone (i.e., the phone number of the store), store pictures, and other information. Correspondingly, the input store data to be processed includes: store address information, store name, store phone, store pictures, and other information of the store to be processed.

[0052] The distance value corresponding to the store to be processed is the distance from the store to be processed to a specific store. The specific store is an existing store selected from each existing store according to a preset rule, and the selection method of the specific store will be introduced in detail below.

[0053] Through the store address information of the store to be processed, the geographical location of the store to be processed, that is, the longitude and latitude information of the store to be processed, can be obtained by using the location interface provided by the map service provider. If the longitude and latitude information cannot be obtained, it is obtained manually on the site of the store to be processed and the corresponding longitude and latitude information is uploaded. In the same way, the longitude and latitude information of the specific store can also be obtained through the store address information of the specific store in the existing store data. The distance from the store to be processed to the specific store can be calculated by using the longitude and latitude information of the store to be processed and the longitude and latitude information of the specific store, so as to obtain the distance value corresponding to the store to be processed.

[0054] Step S202: Search for a target display store in the display store linked list whose difference between the corresponding distance value and the distance value corresponding to the store to be processed is the smallest.

[0055] Among them, the display store is an existing store used for display, and each display store has its own corresponding distance value. The distance value corresponding to the display store is the distance from the display store to a specific store. Using the existing store data (including the display store data), referring to the calculation method of the distance value corresponding to the store to be processed above, the distance value corresponding to the display store can be calculated.

[0056] The display store linked list stores the display store data, and uses a specific store as the first store in the display store linked list. The display store linked list is the main chain of a two-dimensional store linked list. The two-dimensional store linked list also includes sub-chains. Each sub-chain is a duplicate store linked list of a display store, and the duplicate store linked list of a display store stores the duplicate store data of the display store. The process of generating the two-dimensional store linked list is introduced in detail below.

[0057] In the embodiment of the present invention, when calculating the difference between the distance values corresponding to two stores, the larger distance value is subtracted from the smaller distance value so that the calculated difference between the distance values is greater than zero.

[0058] Select a specific store from each existing store according to a preset rule. Specifically, the longitude and latitude information of each existing store can be obtained according to the address information of each store in the existing store data, and according to the longitude and latitude information of each existing store, each existing store is sorted in a predetermined order, and the first existing store after sorting is selected as the specific store. For example: For stores across the country (or the whole province), they are divided into different city subsets according to the city dimension. For stores in the same city subset, through the store address information, use the location interface provided by the map service provider to obtain the geographical location of the store, that is, the longitude and latitude information of the store. If the longitude and latitude information cannot be obtained, it is obtained manually on-site at the store and the corresponding longitude and latitude information is uploaded. They are arranged in ascending order of store longitude. If the longitudes are the same, they are inserted into the queue in ascending order of latitude. If the longitudes and latitudes are both the same, then they are arranged in the order of processing in turn. In this way, a store list is obtained, and the first store in the store list is the specific store. Figure 3 It is a schematic diagram of the store list.

[0059] Calculate the distance values corresponding to each existing store, and sort each existing store in ascending order of the distance value to obtain an existing store linked list. The distance value corresponding to an existing store is the distance from the existing store to the specific store, and this distance is calculated using the longitude and latitude information of the existing store and the specific store. Since the distance value corresponding to the specific store is the distance from it to itself, its corresponding distance value is 0. Therefore, the first existing store in the existing store linked list is the specific store.

[0060] According to the distance values corresponding to each existing store, convert the linked list of existing stores into a two-dimensional linked list. The conversion rule is as follows: Use a specific store as the first existing store in the main chain of the two-dimensional linked list, that is, the initial node of the main chain of the two-dimensional linked list. Then, sequentially determine the subsequent existing stores (i.e., subsequent nodes) in the main chain of the two-dimensional linked list along the linked list of existing stores, and generate sub-chains for each existing store in the main chain of the two-dimensional linked list, such that the difference between the distance values corresponding to adjacent existing stores on the main chain of the two-dimensional linked list is greater than or equal to a preset value, and the difference between the distance value corresponding to an existing store on the main chain of the two-dimensional linked list and the distance value corresponding to its sub-chain existing store is less than this preset value. For example, assume the existing stores are A(0), B(1.1), C(1.2), D(1.4), E(1.5), F(2.1), G(2.5), H(2.8), I(3.6), etc., where the values in parentheses are the distance values corresponding to the existing stores, in kilometers. The sorting of the existing stores in the linked list of existing stores is from A to I, etc., with A as the initial node, and the preset value is set to 1 kilometer (it can be set to other values as needed). Then, starting from node B, calculate the difference between the distance value of B and that of A: 1.1 - 0 = 1.1, which is greater than the preset value of 1. So, node B is the next existing store (the second node) of A on the main chain of the two-dimensional linked list. Continue to calculate the difference between the distance value of C and that of B: 1.2 - 1.1 = 0.1, which is less than the preset value of 1. Then, make C a node on the sub-chain of B, that is, the sub-chain existing store of C. Calculate the differences between the distance values of D and E and that of B in the same way, both of which are less than the preset value of 1. So, D and E are also sub-chain existing stores of C. The sub-chain of B includes three existing stores: C, D, and E. The difference between the distance value of F and that of B is: 2.1 - 1.1 = 1, which is not less than the above preset value. So, F is the third existing store (the third node) on the main chain of the two-dimensional linked list. Similarly, the storage positions of G, H, I, etc. in the two-dimensional linked list can be determined: G and H are sub-chain existing stores of F, and I is the fourth existing store on the main chain of the two-dimensional linked list. In this way, for each existing store, calculate the difference between its distance value and that of the last determined main-chain existing store, and then compare the calculated difference with the preset value, so as to determine whether the existing store is stored in the main chain or the sub-chain of the two-dimensional linked list. In this way, the linked list of existing stores can be converted into a two-dimensional linked list. A schematic diagram of the two-dimensional linked list is as shown in Figure 4 shown below.

[0061] Generate training data according to the sub-chains of the two-dimensional linked list, train a machine learning algorithm model, and obtain the non-repeating stores of all sub-chains of the two-dimensional linked list during the training process. The non-repeating stores of the sub-chain are the stores in the sub-chain that are not repeated with the corresponding existing stores on the main chain of the two-dimensional linked list.

[0062] The machine learning algorithm model in the embodiments of the present invention is a classification-supported model, such as a Bayesian classifier, which can give the similarity probability value of two stores according to the multi-dimensional attribute similarity of the two input stores. Alternatively, the machine learning algorithm model can also adopt other machine learning algorithm models that support classification, such as a support vector machine model, a K-nearest neighbor model (KNN), or a neural network algorithm model.

[0063] The steps of generating training data according to the sub-chains of the two-dimensional linked list, training the machine learning algorithm model, and obtaining the non-duplicate stores of all sub-chains of the two-dimensional linked list during the training process can specifically include:

[0064] According to the calculation rule of the multi-dimensional attribute similarity of stores, calculate the multi-dimensional attribute similarity between the existing stores in all sub-chains and the existing stores in the main chain of the corresponding two-dimensional linked list.

[0065] The multi-dimensional attribute similarity can specifically include distance similarity, phone similarity, name similarity, address similarity, and picture similarity.

[0066] The calculation rule of the multi-dimensional attribute similarity of stores includes calculating the multi-dimensional attribute similarity between the first store and the second store in the following way:

[0067] The distance similarity = 1 - m / k, where m is the difference between the distance values corresponding to the first store and the second store, k is the above preset value, and m / k represents the ratio of m to k;

[0068] For the phone similarity, compare the phone numbers of the first store and the second store digit by digit, count the number of inconsistent digits between the two phone numbers, and find the phone similarity value corresponding to this number of digits. For example, the phone similarity corresponding to two completely identical phone numbers is 1, the phone similarity corresponding to the last digit being different is 0.9, the phone similarity corresponding to the last two digits being different is 0.8, and so on;

[0069] The name similarity and the address similarity are obtained through a text-to-vector model, such as a word2vec model, but not limited to this model;

[0070] The picture similarity is obtained through a picture similarity matching algorithm model. The picture similarity matching algorithm model is, for example, an algorithm model such as a hash algorithm, a convolutional neural network, or a similarity matching algorithm based on local invariant features;

[0071] Among them, the first store is the existing store in the sub-chain of the two-dimensional linked list, and the second store is the existing store in the main chain of the two-dimensional linked list.

[0072] The distance value corresponding to the store is the distance from the store to a specific store. The distance can be calculated using the longitude and latitude information of the store and the specific store, and the longitude and latitude information of the two is obtained according to the address information of the store and the specific store. Information such as the address information, phone number, name, and picture of the store and the specific store can be obtained from the store data and the specific store data accordingly.

[0073] Select the first part of the sub-chains from all the sub-chains. For example, randomly select one-third of the sub-chains as the first part of the sub-chains for manual annotation, that is, identify the stores in the sub-chains that are repeated with the existing stores on the main chain of the two-dimensional linked list (abbreviated as repeated stores). The schematic diagram of the identification of the repeated stores in the sub-chains is as Figure 5 shown, where "√" represents a repeated store and "×" represents a non-repeated store. Generate the first training data according to the first part of the sub-chains and perform the first training on the machine learning algorithm model;

[0074] Select the second part of the sub-chains from the currently remaining sub-chains. For example, select another one-half of the sub-chains from the currently remaining sub-chains as the second part of the sub-chains for testing. Use the machine learning algorithm model trained for the first time to perform similarity prediction on the second part of the sub-chains. The similarity prediction includes predicting the similarity between the existing stores in the sub-chains and the existing stores on the corresponding main chain, and select some sub-chains from the second part of the sub-chains according to the prediction results to generate the second training data. For example, the nodes (existing stores on the sub-chains) with a similarity probability greater than 90% output by the machine learning algorithm model, and use this part of the sub-chain data as the second training data. To ensure the training effect, the second training data can be manually verified and confirmed. Change the ones that are manually confirmed as non-repeated stores to the main chain, and use the existing stores on those sub-chains that are manually confirmed as repeated stores of the existing stores on the main chain as the second training data. Use the first and second training data to perform the second training on the machine learning algorithm model;

[0075] Use the machine learning algorithm model trained for the second time to perform similarity prediction on the last remaining sub-chains, and select some sub-chains from the last remaining sub-chains according to the prediction results to generate the third training data. Referring to the method of generating the second training data, the nodes (existing stores on the sub-chains) with a similarity probability greater than 90% output by the machine learning algorithm model can be selected, and it is manually verified and confirmed whether they are repeated stores in the sub-chains. Finally, use the existing stores on the part of the sub-chains after manual confirmation as the third training data. Use the first, second, and third training data to perform the third training on the machine learning algorithm model; among them, the first, second, and third training data specifically include the multi-dimensional attribute similarity between the existing stores in the sub-chains and the existing stores on the corresponding main chain and the annotation values of the existing stores in the sub-chains, and the annotation values include the marked repeated stores and non-repeated stores in the corresponding sub-chains;

[0076] The input of the above machine learning algorithm model is the multi-dimensional attribute similarity between an existing store in each sub-chain and the existing stores in the main chain of the two-dimensional linked list corresponding to it, and the output is the similarity probability value between the two stores, that is, the prediction result of similarity prediction.

[0077] When the similarity probability output by the machine learning algorithm model is less than or equal to 90% for a node (the existing store on the sub-chain), it is determined as a non-duplicate store of the sub-chain, that is, a store on the sub-chain that is not duplicate with the existing stores in the main chain of the corresponding two-dimensional linked list.

[0078] Through the above process, it can be determined whether the existing stores in each sub-chain are duplicate stores or non-duplicate stores of the existing stores in its main chain. That is, according to the marked non-duplicate stores in each sub-chain and the non-duplicate stores in each sub-chain found by using the prediction results of the machine learning algorithm model, all non-duplicate stores of the sub-chains can be obtained.

[0079] Change the non-duplicate stores of all sub-chains to the main chain of the two-dimensional linked list. The new two-dimensional linked list obtained is the two-dimensional store linked list. The stores of the main chain nodes of the two-dimensional store linked list are all non-duplicate and can be used as the stores to be shown to users (i.e., the display stores). The main chain of the two-dimensional store linked list can be called the display store linked list; the sub-chain nodes of the two-dimensional store linked list are duplicate stores corresponding to the stores of the main chain nodes (display stores). The sub-chain of the two-dimensional store linked list can be called the duplicate store linked list of the display stores. The schematic diagram of the two-dimensional store linked list is as Figure 6 shown, where "√" represents duplicate stores.

[0080] Step S203: Compare the difference between the distance values corresponding to the target display store and the store to be processed with a preset value. If the difference between the distance values is greater than or equal to the preset value, it is determined that there is no duplicate store data of the store to be processed in the display store linked list; if the difference between the distance values is less than the preset value, predict the similarity between the target display store and the store to be processed, and check whether there is duplicate store data of the store to be processed in the display store linked list according to the similarity.

[0081] The steps of predicting the similarity between the target display store and the store to be processed specifically include: calculating the multi-dimensional attribute similarity between the target display store and the store to be processed according to the multi-dimensional attribute similarity calculation rule of the store, and inputting the multi-dimensional attribute similarity into the machine learning algorithm model after the third training (i.e., after all training processes are completed) to predict the similarity between the target display store and the store to be processed.

[0082] Check whether there is duplicate store data of the store to be processed in the display store linked list according to the similarity. Specifically, judge whether the target display store is a duplicate store of the store to be processed according to this similarity. If so, there is a duplicate store of the store to be processed in the display store linked list (this duplicate store is the target display store); if not, there is no duplicate store of the store to be processed in the display store linked list.

[0083] According to the multi-dimensional attribute similarity calculation rule of the store, when calculating the multi-dimensional attribute similarity between the target display store and the store to be processed, refer to the above introduction of calculating the multi-dimensional attribute similarity between the first store and the second store. When calculating the multi-dimensional attribute similarity between the target display store and the store to be processed, the above first store is the store to be processed, and the second store is the target display store.

[0084] After predicting the similarity between the target display store and the store to be processed through the machine learning algorithm model, if the similarity probability output by the machine learning algorithm model is greater than 90%, it means that the model predicts that the target display store is a duplicate store of the store to be processed. In addition, after predicting the duplicate store, manual verification can also be carried out to judge whether the two are really duplicate stores; if the similarity probability is less than or equal to 90%, it means that it is predicted that the target display store is not a duplicate store of the store to be processed.

[0085] If there is no duplicate store of the store to be processed in the display store linked list, store the store data to be processed in the display store linked list for display; if there is a duplicate store of the store to be processed in the display store linked list, store the store data to be processed in the duplicate store linked list of the target display store so as not to be displayed.

[0086] Among them, when storing the store to be processed in the display store linked list, determine the specific storage location of the store to be processed in the display store linked list according to the size of the distance values corresponding to the target display store and the store to be processed. Specifically, if the distance value corresponding to the target display store is greater than the distance value corresponding to the store to be processed, insert a node before the target display store in the display store linked list to store the store data to be processed, otherwise insert a node after the target display store in the display store linked list to store the store data to be processed.

[0087] The embodiment of the present invention also provides a method for processing store off-line, and the main processing steps include:

[0088] After receiving the off-line instruction for a store in the two-dimensional store linked list, find the storage location of the data of the store to be off-line;

[0089] If the data of the store to be taken offline is stored in the displayed store linked list and there is a duplicate store linked list of the store to be taken offline in the two-dimensional store linked list, then delete the stored data of the store to be taken offline, and move the store data with the highest user rating in the duplicate store linked list to the original storage location of the store to be taken offline;

[0090] If the data of the store to be taken offline is stored in the displayed store linked list and there is no duplicate store linked list of the store to be taken offline in the two-dimensional store linked list, or if the data of the store to be taken offline is stored in a duplicate store linked list of a certain displayed store, then directly delete the stored data of the store to be taken offline.

[0091] In addition, in the case of incorrect duplicate store detection due to incorrect store longitude and latitude, it is possible to receive the feedback information that a certain store is a duplicate store input by the user, and verify whether the corresponding store is a duplicate store of other displayed stores through manual verification according to the feedback information. If so, insert the store into the duplicate store linked list of the displayed store to not display it.

[0092] The data processing method of the embodiment of the present invention solves the problem of duplicate display of online stores, so that even if the information provided by the actual same store is different (such as different names), it will not be repeatedly displayed to users in the form of different stores. The embodiment of the present invention can improve the user experience. And it covers all scenarios of store display, including existing store data processing, newly added store data processing, existing store offline processing, etc., realizes the automation of the processing of each store's data, reduces many problems brought by manual review. In addition, by introducing the method of machine learning, as the amount of data increases, the accuracy of review can be significantly improved and the manual workload can be reduced.

[0093] Figure 7 It is a schematic diagram of the main modules of the data processing device according to the embodiment of the present invention.

[0094] As Figure 7 shown, the data processing device 700 of the embodiment of the present invention mainly includes a calculation module 701, a search module 702, and a troubleshooting and processing module 703.

[0095] The calculation module 701 is used to calculate the numerical information corresponding to the entity to be processed according to the input entity data to be processed;

[0096] The search module 702 is used to search for the target displayed entity with the smallest difference between the corresponding numerical information and the numerical information corresponding to the entity to be processed in the displayed entity linked list. The displayed entity linked list stores the displayed entity data. The displayed entity is an existing entity for display, and each displayed entity has its own corresponding numerical information;

[0097] The troubleshooting and processing module 703 is configured to compare the difference between the numerical information corresponding to the target display entity and the to-be-processed entity with a preset value. If the difference in the numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the to-be-processed entity in the display entity linked list. If the difference in the numerical information is less than the preset value, the similarity between the target display entity and the to-be-processed entity is predicted, and whether there is duplicate entity data of the to-be-processed entity in the display entity linked list is checked according to the similarity.

[0098] Among them, the display entity linked list is the main chain of the two-dimensional entity linked list. The two-dimensional entity linked list also includes sub-chains. Each sub-chain is a duplicate entity linked list of a display entity, and the duplicate entity data of the display entity is stored therein.

[0099] The data processing device 700 further includes a two-dimensional linked list generation module, which is configured to obtain the two-dimensional entity linked list in the following manner: select specific entities from each existing entity according to a preset rule; calculate the numerical information corresponding to each existing entity, and sort each existing entity in ascending order according to the numerical information to obtain an existing entity linked list; convert the existing entity linked list into a two-dimensional linked list according to the numerical information corresponding to each existing entity. The conversion rule is: use the specific entity as the first existing entity of the main chain of the two-dimensional linked list, and sequentially determine the subsequent existing entities of the main chain of the two-dimensional linked list along the existing entity linked list, and generate sub-chains for each existing entity of the main chain of the two-dimensional linked list, so that the difference between the numerical information corresponding to adjacent existing entities on the main chain of the two-dimensional linked list is greater than or equal to the preset value, and the difference between the numerical information corresponding to the existing entity on the main chain of the two-dimensional linked list and the existing entity of its sub-chain is less than the preset value; generate training data according to the sub-chains of the two-dimensional linked list, train a machine learning algorithm model, and obtain non-duplicate entities of all sub-chains of the two-dimensional linked list during the training process. The non-duplicate entity of the sub-chain is an entity in the sub-chain that is not duplicate with the existing entity of the corresponding main chain of the two-dimensional linked list. The machine learning algorithm model is a model that supports classification; change the non-duplicate entities of all sub-chains to the main chain of the two-dimensional linked list, and the obtained new two-dimensional linked list is the two-dimensional entity linked list.

[0100] The data processing device 700 may further include a display processing module, which is configured to: if there is no duplicate entity data of the to-be-processed entity in the display entity linked list, store the to-be-processed entity data in the display entity linked list for display; if there is duplicate entity data of the to-be-processed entity in the display entity linked list, store the to-be-processed entity data in the duplicate entity linked list of the target display entity for non-display.

[0101] The data processing device 700 may further include a decommissioning processing module, which is configured to: after receiving a decommissioning instruction for an entity in the two-dimensional entity linked list, find the storage location of the data of the entity to be decommissioned; if the data of the entity to be decommissioned is stored in the display entity linked list and there is a duplicate entity linked list of the entity to be decommissioned in the two-dimensional entity linked list, delete the stored data of the entity to be decommissioned, and move the entity data with the highest user rating in the duplicate entity linked list to the original storage location of the data of the entity to be decommissioned; if the data of the entity to be decommissioned is stored in the display entity linked list and there is no duplicate entity linked list of the entity to be decommissioned in the two-dimensional entity linked list, or if the data of the entity to be decommissioned is stored in the duplicate entity linked list of a certain display entity, directly delete the stored data of the entity to be decommissioned.

[0102] The troubleshooting processing module 703 may include a similarity prediction sub-module, which is configured to: calculate the multi-dimensional attribute similarity between a target display entity and an entity to be processed according to the entity multi-dimensional attribute similarity calculation rule, and input the multi-dimensional attribute similarity into a trained machine learning algorithm model to predict the similarity between the target display entity and the entity to be processed.

[0103] In addition, the specific implementation content of the data processing device in the embodiments of the present invention has been described in detail in the above-described data processing method, so the repeated content will not be described herein.

[0104] Figure 8 An exemplary system architecture 800 to which the data processing method or data processing device of the embodiments of the present invention can be applied is shown.

[0105] As Figure 8 shown, the system architecture 800 may include terminal devices 801, 802, 803, a network 804, and a server 805. The network 804 is used to provide a medium for a communication link between the terminal devices 801, 802, 803 and the server 805. The network 804 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0106] Users may use the terminal devices 801, 802, 803 to interact with the server 805 through the network 804 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 801, 802, 803, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0107] The terminal devices 801, 802, 803 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0108] The server 805 can be a server that provides various services. For example, it can be a background management server (only an example) that supports shopping websites browsed by users using the terminal devices 801, 802, and 803. The background management server can analyze and process data such as product information query requests received, and feedback the processing results (such as target push information - only an example) to the terminal devices.

[0109] It should be noted that the data processing method provided by the embodiments of the present invention is generally executed by the server 805. Correspondingly, the data processing device is generally arranged in the server 805.

[0110] It should be understood that Figure 8 the numbers of the terminal devices, networks, and servers in

[0111] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. Figure 9 Next, refer to Figure 9 which shows a schematic structural diagram of a computer system 900 suitable for implementing the terminal device or server of the embodiments of the present application.

[0112] As Figure 9 shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage section 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0113] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that the computer program read from it can be installed into the storage section 908 as needed.

[0114] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 909 and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-mentioned functions defined in the system of the present application are performed.

[0115] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0117] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a calculation module, a search module, and a troubleshooting processing module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the calculation module can also be described as "a module for calculating the distance value corresponding to the entity to be processed according to the input entity data to be processed".

[0118] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or may exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: calculating numerical information corresponding to the entity to be processed according to the input entity data to be processed; searching in the display entity list for a target display entity with the smallest difference between the corresponding numerical information and the numerical information of the entity to be processed, where the display entity list stores display entity data, the display entity is an existing entity for display, and each display entity has its own corresponding numerical information; comparing the difference between the numerical information of the target display entity and the entity to be processed with a preset value. If the difference in numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the entity to be processed in the display entity list; if the difference in numerical information is less than the preset value, the similarity between the target display entity and the entity to be processed is predicted, and whether there is duplicate entity data of the entity to be processed in the display entity list is checked according to the similarity.

[0119] According to the technical solution of the embodiment of the present invention, calculate the numerical information corresponding to the entity to be processed according to the input entity data to be processed; find the target display entity in the display entity linked list with the smallest difference between the corresponding numerical information and the numerical information corresponding to the entity to be processed; compare the difference between the numerical information corresponding to the target display entity and the entity to be processed with a preset value. If the difference in numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the entity to be processed in the display entity linked list; if the difference in numerical information is less than the preset value, predict the similarity between the target display entity and the entity to be processed, and check whether there is duplicate entity data of the entity to be processed in the display entity linked list according to the similarity. This implementation method can improve data review efficiency, reduce the review workload, shorten the review time, reduce the dependence on manual experience, reduce review errors, and improve the user experience.

[0120] The above specific implementation manners do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data processing method, characterized in that, it includes: calculating numerical information corresponding to the entity to be processed according to the input entity data to be processed; finding a target display entity in the display entity linked list with the smallest difference between the corresponding numerical information and the numerical information corresponding to the entity to be processed, where the display entity linked list stores display entity data, the display entity is an existing entity for display, and each display entity has its own corresponding numerical information; comparing the difference between the numerical information corresponding to the target display entity and the entity to be processed with a preset value. If the difference in numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the entity to be processed in the display entity linked list; if the difference in numerical information is less than the preset value, predicting the similarity between the target display entity and the entity to be processed, and checking whether there is duplicate entity data of the entity to be processed in the display entity linked list according to the similarity.

2. The method according to claim 1, characterized in that, the display entity linked list is the main chain of a two-dimensional entity linked list, and the two-dimensional entity linked list further includes sub-chains. Each sub-chain is a duplicate entity linked list of a display entity, which stores the duplicate entity data of the display entity. The two-dimensional entity linked list is obtained in the following manner: selecting specific entities from each existing entity according to a preset rule; calculating the numerical information corresponding to each existing entity, and sorting each existing entity in ascending order of the numerical information to obtain an existing entity linked list; converting the existing entity linked list into a two-dimensional linked list according to the numerical information corresponding to each existing entity. The conversion rule is: using the specific entity as the first existing entity of the main chain of the two-dimensional linked list, and sequentially determining the subsequent existing entities of the main chain of the two-dimensional linked list along the existing entity linked list, and generating sub-chains for each existing entity of the main chain of the two-dimensional linked list, so that the difference between the numerical information corresponding to adjacent existing entities on the main chain of the two-dimensional linked list is greater than or equal to the preset value, and the difference between the numerical information corresponding to the existing entity on the main chain of the two-dimensional linked list and the existing entity of its sub-chain is less than the preset value; generating training data according to the sub-chains of the two-dimensional linked list, training a machine learning algorithm model, and obtaining non-duplicate entities of all sub-chains of the two-dimensional linked list during the training process. The non-duplicate entity of the sub-chain is an entity in the sub-chain that is not duplicate with the existing entity on the corresponding main chain of the two-dimensional linked list. The machine learning algorithm model is a model for supporting classification; changing the non-duplicate entities of all sub-chains to the main chain of the two-dimensional linked list, and the obtained new two-dimensional linked list is the two-dimensional entity linked list.

3. The method according to claim 2, characterized in that, it further includes: if there is no duplicate entity data of the entity to be processed in the display entity linked list, storing the entity data to be processed in the display entity linked list for display; if there is duplicate entity data of the entity to be processed in the display entity linked list, storing the entity data to be processed in the duplicate entity linked list of the target display entity for non-display.

4. The method according to claim 2, characterized in that, the method further includes: After receiving a logout instruction for an entity in the two-dimensional entity linked list, find the storage location of the data of the entity to be logged out; If the data of the entity to be logged out is stored in the display entity linked list and there is a duplicate entity linked list of the entity to be logged out in the two-dimensional entity linked list, delete the stored data of the entity to be logged out, and move the entity data with the highest user rating in the duplicate entity linked list to the original storage location of the data of the entity to be logged out; If the data of the entity to be logged out is stored in the display entity linked list and there is no duplicate entity linked list of the entity to be logged out in the two-dimensional entity linked list, or if the data of the entity to be logged out is stored in the duplicate entity linked list of a certain display entity, directly delete the stored data of the entity to be logged out.

5. The method according to claim 2, wherein, The steps of generating training data according to the sub-chains of the two-dimensional linked list, training a machine learning algorithm model, and obtaining non-duplicate entities of all sub-chains of the two-dimensional linked list during the training process include: Calculate the multi-dimensional attribute similarity between the existing entities in all the sub-chains and the existing entities in the main chain of the corresponding two-dimensional linked list according to the multi-dimensional attribute similarity calculation rule of the entity; Select a first part of sub-chains from all the sub-chains, generate first training data according to the first part of sub-chains, and perform first training on the machine learning algorithm model; Select a second part of sub-chains from the currently remaining sub-chains, perform similarity prediction on the second part of sub-chains by using the machine learning algorithm model trained in the first training, the similarity prediction includes predicting the similarity between the existing entities in the sub-chain and the existing entities in the corresponding main chain, and select some sub-chains from the second part of sub-chains according to the prediction result to generate second training data, and perform second training on the machine learning algorithm model by using the first and second training data; Perform the similarity prediction on the last remaining sub-chains by using the machine learning algorithm model trained in the second training, and select some sub-chains from the last remaining sub-chains according to the prediction result to generate third training data, and perform third training on the machine learning algorithm model by using the first, second and third training data; wherein, the first, second and third training data respectively include the multi-dimensional attribute similarity and annotation value between the existing entities in the corresponding sub-chains and the existing entities in the corresponding main chain, and the annotation value includes the non-duplicate entities of the corresponding sub-chains marked; Obtain the non-duplicate entities of all the sub-chains according to the non-duplicate entities of each sub-chain marked and the non-duplicate entities in each sub-chain found by using each prediction result of the machine learning algorithm model.

6. The method according to claim 5, wherein, The steps of predicting the similarity between the target display entity and the entity to be processed include: Calculate the multi-dimensional attribute similarity between the target display entity and the entity to be processed according to the multi-dimensional attribute similarity calculation rule of the entity, and input the multi-dimensional attribute similarity into the machine learning algorithm model trained in the third training to predict the similarity between the target display entity and the entity to be processed.

7. The method according to claim 6, It is characterized in that the multi-dimensional attribute similarity includes distance similarity, phone similarity, name similarity, address similarity, and picture similarity, and the calculation rule of the entity multi-dimensional attribute similarity includes calculating the multi-dimensional attribute similarity between the first entity and the second entity in the following way: The distance similarity = 1 - m / k, where m is the difference between the numerical information corresponding to the first entity and the second entity, k is the preset value, m / k represents the ratio of m to k, the numerical information corresponding to the first entity is the distance from the first entity to the specific entity, which is calculated using the longitude and latitude information of the first entity and the specific entity; the numerical information corresponding to the second entity is the distance from the second entity to the specific entity, which is calculated using the longitude and latitude information of the second entity and the specific entity; The phone similarity is obtained by comparing each digit of the phone numbers of the first entity and the second entity, counting the number of inconsistent digits between the two phone numbers, and looking up the phone similarity value corresponding to the number of digits; The name similarity and the address similarity are obtained through a text-to-vector model; The picture similarity is obtained through a picture similarity matching algorithm model; where the first entity is the existing entity of the two-dimensional linked list sub-chain, the second entity is the existing entity of the two-dimensional linked list main chain, or the first entity is the entity to be processed, and the second entity is the target display entity.

8. A data processing device It is characterized in that comprising: a calculation module for calculating the numerical information corresponding to the entity to be processed according to the input data of the entity to be processed; a search module for searching in the display entity linked list for the target display entity with the smallest difference between the corresponding numerical information and the numerical information corresponding to the entity to be processed, the display entity linked list storing display entity data, the display entity being an existing entity for display, and each display entity having its own corresponding numerical information; a troubleshooting and processing module for comparing the difference between the numerical information corresponding to the target display entity and the entity to be processed with a preset value. If the difference in numerical information is greater than or equal to the preset value, it is determined that there is no duplicate entity data of the entity to be processed in the display entity linked list; if the difference in numerical information is less than the preset value, the similarity between the target display entity and the entity to be processed is predicted, and whether there is duplicate entity data of the entity to be processed in the display entity linked list is checked according to the similarity.

9. An electronic device It is characterized in that comprising: one or more processors; a memory for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1 - 7.

10. A computer-readable medium having a computer program stored thereon, It is characterized in that when the program is executed by a processor, it implements the method according to any one of claims 1 - 7.

Citation Information

Patent Citations

  • Information processing method and device

    CN103198004A

  • Data processing method, device and system

    CN103970795A