Data processing method and device and electronic equipment
By calculating the similarity between products and performing deduplication operations, the problem of high repetition rate of goods in the home improvement industry is solved, and standardized management of goods and inventory optimization are achieved.
Patent Information
- Application Number
- CN202510293393.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-08
AI Technical Summary
In the home decoration industry, there are many types of products, diverse brands, and different specifications, resulting in a large number of similar or even duplicate products in the product creation process, affecting the standardized management, inventory management and business statistical analysis of products.
By obtaining the product information to be deduplicated, the similarity between the products is calculated, the target product cluster is determined, and the deduplication operation is performed based on the similarity detection results to delete duplicate products.
It reduces the repetition rate of goods in the database and improves the standardized management and inventory management efficiency of goods.
Smart Images

Figure CN120277062A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a data processing method, apparatus, and electronic device. Background Art
[0002] Currently, in the home improvement industry, there are a wide variety of products, diverse brands, and different specifications. This diversity can easily lead to the appearance of a large number of similar or even duplicate products during the product creation process, thereby affecting the standardized management, inventory management, and business statistical analysis of products.
[0003] Therefore, how to reduce the duplication rate of products in the database has become an urgent problem to be solved. Summary of the Invention
[0004] To solve the above technical problems, the present disclosure provides a data processing method, apparatus, and electronic device.
[0005] In a first aspect, the present disclosure provides a data processing method, including: obtaining product information of at least one actual product to be de-duplicated; where the product information includes one or more of product category, product brand, product specification, and product model; calculating the similarity between different actual products based on the product information; determining at least one group of target products in each cluster based on the similarity; where the target products include any one of the actual products; performing similarity detection on the product information corresponding to each group of target products to obtain a duplicate detection result; and performing a de-duplication operation on the actual products based on the duplicate detection result corresponding to each cluster to obtain the de-duplicated actual products.
[0006] In a second aspect, the present disclosure provides a data processing apparatus, including: an obtaining unit for obtaining product information of at least one actual product to be de-duplicated; where the product information includes one or more of product category, product brand, product specification, and product model; a processing unit for calculating the similarity between different actual products based on the product information obtained by the obtaining unit; the processing unit is further configured to determine at least one group of target products in each cluster based on the similarity; where the target products include any one of the actual products; the processing unit is further configured to perform similarity detection on the product information corresponding to each group of target products obtained by the obtaining unit to obtain a duplicate detection result; and the processing unit is further configured to perform a de-duplication operation on the actual products based on the duplicate detection result corresponding to each cluster to obtain the de-duplicated actual products.
[0007] In a third aspect, the present invention provides an electronic device, including: a memory and a processor, the memory is used to store a computer program; the processor is configured to cause the electronic device to implement the data processing method according to any one of the first aspect when executing the computer program.
[0008] Fourth aspect, the present invention provides a computer-readable storage medium, including: a computer program stored on the computer-readable storage medium, and the computer program is executed by a controller to perform the data processing method according to any one of the first aspect.
[0009] Fifth aspect, the present invention provides a computer program product, when the computer program product runs on a computer, it causes the computer to execute the data processing method according to any one of the first aspect.
[0010] These aspects or other aspects of the present disclosure will be more clearly understood in the following description.
[0011] The technical solution provided by the present disclosure has the following advantages compared with the prior art:
[0012] For the data processing method provided by the present disclosure, by obtaining the product information of at least one actual product to be de-duplicated; then, based on the product information, calculating the similarity between different actual products; based on the similarity, determining at least one group of target products in each cluster; performing similarity detection based on the product information corresponding to each group of target products to obtain a duplicate detection result; thus, based on the duplicate detection result, it can be determined whether there are duplicate products in each group of target products. For example, when the duplicate detection result is that the products are duplicates, any one of the actual products with the duplicate detection result of product duplication can be deleted to obtain the de-duplicated actual products. Then, based on the duplicate detection result corresponding to each cluster, perform a de-duplication operation on the actual products to obtain the de-duplicated actual products, thus completing the de-duplication of the actual products that need to be de-duplicated. Since the duplicate products in the database are reduced, the duplication rate of the products in the database can be reduced, and the problem of how to reduce the duplication rate of the products in the database is solved. Description of the Drawings
[0013] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0015] Figure 1 Exemplarily shows one of the flow diagrams of a data processing method provided in Embodiment 1;
[0016] Figure 2 Exemplarily shows the second flow diagram of a data processing method provided in Embodiment 1;
[0017] Figure 3 Exemplarily shown is the third flowchart of a data processing method provided in the first embodiment of the present disclosure;
[0018] Figure 4 Exemplarily shown is the fourth flowchart of a data processing method provided in the first embodiment of the present disclosure;
[0019] Figure 5 Exemplarily shown is the fifth flowchart of a data processing method provided in the first embodiment of the present disclosure;
[0020] Figure 6 Exemplarily shown is the sixth flowchart of a data processing method provided in the first embodiment of the present disclosure;
[0021] Figure 7 Exemplarily shown is the seventh flowchart of a data processing method provided in the first embodiment of the present disclosure;
[0022] Figure 8 Exemplarily shown is the structural diagram of a data processing apparatus provided in the second embodiment of the present disclosure;
[0023] Figure 9 Exemplarily shown is the structural diagram of an electronic device provided in the third embodiment of the present disclosure. Detailed implementation manners
[0024] In order to more clearly understand the above-mentioned objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0025] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0026] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0027] In some examples, ChatGPT in the embodiments of the present disclosure refers to a large language model developed by OpenAI.
[0028] Example 1
[0029] Figure 1 A schematic flowchart of a data processing method is exemplarily shown. The execution subject of this example can be a server, such as Figure 1 As shown, the method includes:
[0030] S11. Obtain the product information of at least one actual product to be de-duplicated. Among them, the product information includes one or more of product category, product brand, product specification and product model.
[0031] In some examples, the product specification includes dimension information, color information, etc. The dimension information includes length, width and height, or radius, thickness, transparency, etc.
[0032] S12. Calculate the similarity between different actual products based on the product information.
[0033] In some examples, calculating the similarity between different actual products based on the product information includes: calculating the similarity between each actual product and other actual products except this actual product. For example, if the at least one actual product to be de-duplicated includes actual product 1, actual product 2 and actual product 3, then calculating the similarity between different actual products based on the product information includes calculating the similarity between actual product 1 and actual product 2, calculating the similarity between actual product 1 and actual product 3, and calculating the similarity between actual product 2 and actual product 3.
[0034] In some examples, when calculating the similarity between different actual products based on product information, it is necessary to calculate the similarity between each actual product and other actual products except this actual product. When the number of actual products is large, it will consume a large amount of computing resources. Therefore, the actual products can be clustered first to obtain at least one cluster. For example, clustering the actual products based on product information to obtain at least one cluster. After that, calculate the similarity between each actual product in each cluster and other actual products except this actual product, which can greatly reduce the consumption of computing resources and improve the user experience.
[0035] In some examples, when calculating the similarity between different actual products based on product information, since the product information may contain some dirty data (such as &, * etc.) and the data formats are not unified (such as traditional Chinese to simplified Chinese, capital letters to lowercase etc.), it is necessary to preprocess the product information first to generate preset information. Among them, the preprocessing includes one or more of stop word removal operations, filtering operations, and text format conversion operations to ensure the uniformity of the data. After that, feature extraction is performed on the preset information to obtain the feature vector corresponding to the preset information. Based on the feature vectors corresponding to each actual product, the similarity between different actual products is determined.
[0036] In some examples, the preset information can be input into a pre-trained language model (Bidirectional Encoder Representations from Transformers, BERT) for vector conversion to obtain the feature vector corresponding to the preset information.
[0037] In some examples, a product identifier can be set for each actual product, such as an identification code ID. After that, establish the corresponding relationship between the identification code ID and the feature vector corresponding to the actual product, and store this corresponding relationship in the memory of the server. In this way, when it is necessary to deduplicate the actual products, if it is determined based on the identification code ID of this actual product that there is an ID in the database that is the same as this identification code ID, then directly read the feature vector corresponding to this identification code ID, without having to calculate the feature vector of this actual product again, reducing the consumption of computing resources and shortening the data processing time.
[0038] In some examples, when calculating the similarity between different actual products based on product information, the product information can be preprocessed first to generate preset information. Among them, the preprocessing includes one or more of stop word removal operations, filtering operations, and text format conversion operations, so as to ensure the unity of data. After that, feature extraction is performed on the preset information to obtain the feature vectors corresponding to the preset information. Then, the feature vectors are aggregated to obtain at least one cluster. Finally, the similarity between the feature vectors of each actual product in each cluster and the feature vectors of other actual products except this actual product is calculated, which can greatly reduce the occupation of computing resources and improve the user experience.
[0039] In some examples, the similarity between different actual products is equal to the cosine similarity between the feature vectors of different actual products. For example, if at least one actual product to be de-duplicated includes actual product 1, actual product 2, and actual product 3, then based on the product information, calculating the similarity between different actual products includes calculating the cosine similarity between actual product 1 and actual product 2, calculating the cosine similarity between actual product 1 and actual product 3, and calculating the cosine similarity between actual product 2 and actual product 3. After that, the cosine similarity between actual product 1 and actual product 2 is used as the similarity between actual product 1 and actual product 2, the cosine similarity between actual product 1 and actual product 3 is used as the similarity between actual product 1 and actual product 3, and the cosine similarity between actual product 2 and actual product 3 is used as the similarity between actual product 2 and actual product 3.
[0040] Alternatively, the similarity between different actual products is equal to the target distance between the feature vectors of different actual products. For example, if at least one actual product to be de-duplicated includes actual product 1, actual product 2, and actual product 3, then based on the product information, calculating the similarity between different actual products includes calculating the target distance between actual product 1 and actual product 2, calculating the target distance between actual product 1 and actual product 3, and calculating the target distance between actual product 2 and actual product 3. After that, the target distance between actual product 1 and actual product 2 is used as the similarity between actual product 1 and actual product 2, the target distance between actual product 1 and actual product 3 is used as the similarity between actual product 1 and actual product 3, and the target distance between actual product 2 and actual product 3 is used as the similarity between actual product 2 and actual product 3.
[0041] In some examples, the product information can be input into a similarity model for similarity calculation to obtain the similarity between different actual products. Among them, the training process of the similarity model includes:
[0042] Obtain the first training sample data and the first labeling result of the first training sample data. Among them, the first training sample data includes at least one set of historical commodity information, and the first labeling result includes the similarity between different commodities in each set of historical commodity information.
[0043] Input the first training sample data into the first neural network model for learning to obtain the first prediction result of the first neural network model for the first training sample data.
[0044] Based on the first prediction result and the first labeling result, adjust the network parameters of the first neural network model until the first neural network model converges to obtain a similarity model.
[0045] S13. Determine at least one set of target commodities in each cluster based on the similarity; among them, the target commodities include any one of the actual commodities.
[0046] In some examples, the target commodity is an actual commodity with a similarity greater than the similarity threshold, that is, the actual commodities with a similarity greater than the similarity threshold in the same cluster are target commodities. For example: combining the example given in S12 above, when the similarity between actual commodity 1 and actual commodity 2 is greater than the similarity threshold, at this time, determine actual commodity 1 and actual commodity 2 as a set of target commodities.
[0047] S14. Perform similarity detection based on the commodity information corresponding to each set of target commodities to obtain a duplicate detection result.
[0048] In some examples, although the similarity between the two actual commodities corresponding to each set of target commodities is relatively high, it cannot be directly determined that the two actual commodities are the same product. Therefore, it is necessary to further perform similarity detection based on the commodity information corresponding to each set of target commodities to determine the duplicate detection result. For example, when the duplicate detection result is that the commodities are duplicates, it means that the actual commodities are the same product. Therefore, any one of the two actual commodities can be deleted, so that duplicate products can be removed. For example: combining the example given in S13 above, if the similarity detection is performed on the commodity information corresponding to actual commodity 1 and actual commodity 2, and the duplicate detection result is that the commodities are duplicates, then actual commodity 1 and actual commodity 2 above are the same product. Therefore, any one of actual commodity 1 and actual commodity 2 can be deleted, so that the de-duplication operation of the actual commodities can be completed.
[0049] In some examples, the product information further includes view information, which includes a front view, a left view, and a top view. Then, the server constructs products for each group of target products based on the product specifications and view information in the product information, and obtains the product images of each actual product in each group of target products. Then, by performing similarity detection on the product images of the actual products in each group of target products, a duplicate detection result is obtained. For example, by calculating the image similarity of the product images of the actual products in each group of target products and obtaining each group of duplicate detection results based on the image similarity, if the image similarity is greater than the duplicate threshold, it is determined that the duplicate detection result is that the products are duplicates.
[0050] S15. Based on the duplicate detection result corresponding to each cluster, perform a deduplication operation on the actual products to obtain the deduplicated actual products.
[0051] As can be seen from the above, the data processing method provided by the embodiments of the present disclosure includes: obtaining the product information of at least one actual product to be deduplicated; then, based on the product information, calculating the similarity between different actual products; based on the similarity, determining at least one group of target products in each cluster; performing similarity detection based on the product information corresponding to each group of target products to obtain a duplicate detection result; in this way, it is possible to determine whether there are duplicate products in each group of target products based on the duplicate detection result. For example, when the duplicate detection result is that the products are duplicates, any one of the actual products with the duplicate detection result of product duplication can be deleted to obtain the deduplicated actual products. Then, based on the duplicate detection result corresponding to each cluster, perform a deduplication operation on the actual products to obtain the deduplicated actual products. In this way, the deduplication of the actual products that need to be deduplicated is completed. Since the duplicate products in the database are reduced, the duplication rate of the products in the database can be reduced.
[0052] In some feasible examples, in combination with Figure 1 , such as Figure 2 shown, the above S12 can be specifically implemented by the following S120 and S121.
[0053] S120. Preprocess the product information to generate preset information. Among them, the preprocessing includes one or more of stop word removal operations, filtering operations, and text format conversion operations.
[0054] S121. Extract features from the preset information to obtain the feature vector corresponding to the preset information.
[0055] S122. Based on the feature vector corresponding to each actual product, determine the similarity between different actual products.
[0056] As can be seen from the above, the data processing method provided by the embodiments of the present disclosure includes: obtaining the product information of at least one actual product to be de-duplicated; then preprocessing the product information to generate preset information; extracting features from the preset information to obtain a feature vector corresponding to the preset information; determining the similarity between different actual products based on the feature vector corresponding to each actual product; determining at least one group of target products in each cluster based on the similarity; performing a similarity detection based on the product information corresponding to each group of target products to obtain a duplicate detection result; thus, it is possible to determine whether there are duplicate products in each group of target products based on the duplicate detection result. For example, when the duplicate detection result indicates product duplication, any one of the actual products with a duplicate detection result of product duplication can be deleted to obtain the de-duplicated actual products. Then, based on the duplicate detection result corresponding to each cluster, the actual products are de-duplicated to obtain the de-duplicated actual products, thus completing the de-duplication of the actual products that need to be de-duplicated. Since the number of duplicate products in the database is reduced, the duplication rate of the products in the database can be lowered.
[0057] In some feasible examples, in combination with Figure 2 , such as Figure 3 shown, the above S122 can be specifically implemented by the following S1220 and S1221.
[0058] S1220: Determine the cosine similarity between the feature vectors of different actual products based on the feature vector corresponding to each actual product;
[0059] S1221: Use the cosine similarity as the similarity between the feature vectors of different actual products.
[0060] As can be seen from the above, the data processing method provided by the embodiments of the present disclosure includes: obtaining the product information of at least one actual product to be de-duplicated; then calculating the similarity between different actual products based on the product information; determining at least one group of target products in each cluster based on the similarity; performing a similarity detection based on the product information corresponding to each group of target products to obtain a duplicate detection result; thus, it is possible to determine whether there are duplicate products in each group of target products based on the duplicate detection result. For example, when the duplicate detection result indicates product duplication, any one of the actual products with a duplicate detection result of product duplication can be deleted to obtain the de-duplicated actual products. Then, based on the duplicate detection result corresponding to each cluster, the actual products are de-duplicated to obtain the de-duplicated actual products, thus completing the de-duplication of the actual products that need to be de-duplicated. Since the number of duplicate products in the database is reduced, the duplication rate of the products in the database can be lowered.
[0061] In some feasible examples, in combination with Figure 2 , such as Figure 4As shown, the above S122 can be specifically implemented by the following S1222 and S1223.
[0062] S1222: Based on the feature vectors corresponding to each actual commodity, determine the target distance between the feature vectors of different actual commodities; wherein, the target distance includes any one of the Euclidean distance and the Manhattan distance.
[0063] S1223: Use the target distance as the similarity between the feature vectors of different actual commodities.
[0064] As can be seen from the above, the data processing method provided by the embodiments of the present disclosure includes: obtaining the commodity information of at least one actual commodity to be de-duplicated; then, based on the feature vectors corresponding to each actual commodity, determining the target distance between the feature vectors of different actual commodities; using the target distance as the similarity between the feature vectors of different actual commodities; based on the similarity, determining at least one group of target commodities in each cluster; performing similarity detection based on the commodity information corresponding to each group of target commodities to obtain a duplicate detection result; in this way, it is possible to determine whether there are duplicate products among the commodities in each group of target commodities based on the duplicate detection result. For example, when the duplicate detection result is that there are duplicates, any one of the actual commodities with the duplicate detection result of commodity duplication can be deleted to obtain the de-duplicated actual commodities. Then, based on the duplicate detection result corresponding to each cluster, perform a de-duplication operation on the actual commodities to obtain the de-duplicated actual commodities, thus completing the de-duplication of the actual commodities that need to be de-duplicated. Since the number of duplicate commodities in the database is reduced, the duplication rate of the commodities in the database can be lowered.
[0065] In some feasible examples, in combination with Figure 1 , such as Figure 5 As shown, the above S13 can be specifically implemented by the following S130.
[0066] S130: Use the actual commodities in each cluster with a similarity greater than the similarity threshold as a group of target commodities.
[0067] As can be seen from the above, the data processing method provided by the embodiments of the present disclosure includes: obtaining the product information of at least one actual product to be de-duplicated; then, based on the product information, calculating the similarity between different actual products; taking the actual products with similarity greater than the similarity threshold in each cluster as a group of target products; performing similarity detection based on the product information corresponding to each group of target products to obtain a duplicate detection result; in this way, based on the duplicate detection result, it can be determined whether there are duplicate products in each group of target products. For example, when the duplicate detection result indicates product duplication, any one of the actual products with the duplicate detection result of product duplication can be deleted to obtain the de-duplicated actual products. Then, based on the duplicate detection result corresponding to each cluster, de-duplication operations are performed on the actual products to obtain the de-duplicated actual products, thus completing the de-duplication of the actual products that need to be de-duplicated. Since the number of duplicate products in the database is reduced, the duplication rate of the products in the database can be lowered.
[0068] In some feasible examples, in combination with Figure 1 , such as Figure 6 shown, the above S14 can be specifically implemented through the following S140.
[0069] S140: Input the product information of each group of target products into a detection model for similarity detection to obtain a duplicate detection result; wherein, the detection model is trained based on the product information of historical similar products.
[0070] In some examples, the training process of the detection model includes:
[0071] Obtaining second training sample data and second labeling results of the second training sample data; wherein, the second training sample data includes the product information of at least one group of historical target products, and the second labeling results include the duplicate detection results of each group of target products, and the duplicate detection results include product duplication and non-duplication of products.
[0072] Inputting the second training sample data into a second preset model for learning to obtain a second prediction result of the second preset model for the second training sample data.
[0073] Based on the second prediction result and the second labeling result, adjusting the network parameters of the second preset model until the second preset model converges to obtain the detection model.
[0074] In some examples, the second preset model can be a neural network model or a language model. For example, the language model can be ChatGPT.
[0075] As described above, the data processing method provided by the embodiments of the present disclosure includes: obtaining the product information of at least one actual product to be de-duplicated; then, based on the product information, calculating the similarity between different actual products; based on the similarity, determining at least one group of target products in each cluster; inputting the product information of each group of target products into a detection model for similarity detection to obtain a duplicate detection result; in this way, it is possible to determine whether there are duplicate products in each group of target products based on the duplicate detection result. For example, when the duplicate detection result indicates product duplication, any one of the actual products with a duplicate detection result of product duplication can be deleted to obtain the de-duplicated actual products. Then, based on the duplicate detection result corresponding to each cluster, the de-duplication operation is performed on the actual products to obtain the de-duplicated actual products, thus completing the de-duplication of the actual products that need to be de-duplicated. Since the number of duplicate products in the database is reduced, the duplicate rate of the products in the database can be lowered.
[0076] In some feasible examples, the duplicate detection result includes: product duplication; combined with Figure 1 , such as Figure 7 shown, the above S15 can be specifically implemented by the following S150.
[0077] S150. Based on the duplicate detection result corresponding to each cluster, delete any one of the actual products with a duplicate detection result of product duplication to obtain the de-duplicated actual products.
[0078] In some examples, combined with the examples given in the above S11 and S13, assume that actual product 1 and actual product 2 are a group of target products, and the duplicate detection result of actual product 1 and actual product 2 is product duplication. Thus, any one of actual product 1 and actual product 2 can be deleted. For example, if actual product 1 is deleted, the de-duplicated actual products at this time include actual product 2 and actual product 3.
[0079] Or, assume that actual product 1 and actual product 2 are a group of target products, and the duplicate detection result of actual product 1 and actual product 2 is product non-duplication. At this time, no deletion is performed on actual product 1 and actual product 2.
[0080] As described above, the data processing method provided by the embodiments of the present disclosure includes: obtaining product information of at least one actual product to be de-duplicated; then, based on the product information, calculating the similarity between different actual products; based on the similarity, determining at least one group of target products in each cluster; performing similarity detection based on the product information corresponding to each group of target products to obtain a duplicate detection result; in this way, it is possible to determine whether there are duplicate products among the products in each group of target products based on the duplicate detection result. For example, when the duplicate detection result indicates product duplication, any one of the actual products with a duplicate detection result of product duplication can be deleted to obtain the de-duplicated actual products. Then, based on the duplicate detection result corresponding to each cluster, de-duplication operations are performed on the actual products to obtain the de-duplicated actual products, thus completing the de-duplication of the actual products that need to be de-duplicated. Since the number of duplicate products in the database is reduced, the duplication rate of the products in the database can be lowered.
[0081] Embodiment 2
[0082] The structural schematic diagram of the data processing device provided by the second embodiment of the present application is shown as Figure 8 shown. The data processing device includes: an obtaining unit 201 and a processing unit 202.
[0083] The obtaining unit 201 is configured to obtain product information of at least one actual product to be de-duplicated; wherein the product information includes one or more of product category, product brand, product specification, and product model.
[0084] The processing unit 202 is configured to calculate the similarity between different actual products based on the product information obtained by the obtaining unit 201.
[0085] The processing unit 202 is further configured to determine at least one group of target products in each cluster based on the similarity; wherein the target products include any one of the actual products.
[0086] The processing unit 202 is further configured to perform similarity detection based on the product information corresponding to each group of target products obtained by the obtaining unit 201 to obtain a duplicate detection result.
[0087] The processing unit 202 is further configured to perform de-duplication operations on the actual products based on the duplicate detection result corresponding to each cluster to obtain the de-duplicated actual products.
[0088] In some feasible examples, the processing unit 202 is specifically configured to preprocess the product information obtained by the obtaining unit 201 to generate preset information; wherein, the preprocessing includes one or more of stop word removal operations, filtering operations, and text format conversion operations; the processing unit 202 is specifically configured to extract features from the preset information to obtain a feature vector corresponding to the preset information; the processing unit 202 is specifically configured to determine the similarity between different actual products based on the feature vectors corresponding to each actual product.
[0089] In some feasible examples, the processing unit 202 is specifically configured to determine the cosine similarity between the feature vectors of different actual products based on the feature vectors corresponding to each actual product; the processing unit 202 is specifically configured to use the cosine similarity as the similarity between the feature vectors of different actual products.
[0090] In some feasible examples, the processing unit 202 is specifically configured to determine the target distance between the feature vectors of different actual products based on the feature vectors corresponding to each actual product; wherein, the target distance includes any one of Euclidean distance and Manhattan distance; the processing unit 202 is specifically configured to use the target distance as the similarity between the feature vectors of different actual products.
[0091] In some feasible examples, the processing unit 202 is specifically configured to use the actual products in each cluster with a similarity greater than the similarity threshold as a group of target products.
[0092] In some feasible examples, the processing unit 202 is specifically configured to input the product information of each group of target products into a detection model for similarity detection to obtain a duplicate detection result; wherein, the detection model is trained based on the product information of historical similar products.
[0093] In some feasible examples, the duplicate detection result includes: product duplication; the processing unit 202 is specifically configured to delete any actual product with a duplicate detection result of product duplication based on the duplicate detection result corresponding to each cluster to obtain the de-duplicated actual products.
[0094] Wherein, all relevant contents of each step involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and their functions will not be elaborated here.
[0095] Of course, the data processing device provided in the embodiments of the present invention includes but is not limited to the above modules. For example, the data processing device may further include a storage unit 203. The storage unit 203 can be used to store the program code of the data processing device and can also be used to store the data generated during the operation of the data processing device, such as diagnostic data, etc.
[0096] Embodiment III
[0097] Embodiment 3 of the present invention provides a schematic structural diagram of an electronic device. As shown Figure 9 the electronic device may include: at least one processor 51, a memory 52, a communication interface 53, and a communication bus 54.
[0098] The following is a specific introduction to each component of the electronic device:
[0099] Among them, the processor 51 is the control center of the electronic device, which may be a single processor or a collective term for multiple processing elements. For example, the processor 51 is a central processing unit (CPU), or may be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more DSPs, or one or more field programmable gate arrays (FPGAs).
[0100] In a specific implementation, as an embodiment, the processor 51 may include one or more CPUs. For example, the CPUs include CPU0 and CPU1. And, as an embodiment, the electronic device may include multiple processors. For example, the CPUs include processor 51 and processor 55. Each of these processors may be a single-core processor (Single-CPU) or a multi-core processor (Multi-CPU). Here, the processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0101] The memory 52 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 52 can exist independently and be connected to the processor 51 through the communication bus 54. The memory 52 can also be integrated with the processor 51.
[0102] In a specific implementation, the memory 52 is used to store the data in the present invention and execute the software program of the present invention. The processor 51 can execute various functions of the air conditioner by running or executing the software program stored in the memory 52 and calling the data stored in the memory 52.
[0103] The communication interface 53 uses any device such as a transceiver to communicate with other devices or communication networks, such as a radio access network (RAN), a wireless local area network (WLAN), a terminal, the cloud, etc. The communication interface 53 can include an acquisition unit to implement the acquisition function.
[0104] The communication bus 54 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used to represent it, but it does not mean that there is only one bus or one type of bus.
[0105] As an example, in combination with Figure 8, the function implemented by the acquisition unit 201 of the data processing device is the same as that of the communication interface 53, the function implemented by the processing unit 202 in the data processing device is the same as that of the processor 51, and the function implemented by the storage unit 203 in the data processing device is the same as that of the memory 52.
[0106] Embodiment 4
[0107] Embodiment 4 of the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method in any one of the embodiments.
[0108] Embodiment 5
[0109] Embodiment 5 of the present invention provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the method in any one of the embodiments.
[0110] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data processing method, characterized in that, Including: Obtain the product information of at least one actual product to be deduplicated; wherein, the product information includes one or more of product category, product brand, product specification, and product model; Calculate the similarity between different actual products based on the product information; Based on the similarity, determine at least one group of target products in each cluster; wherein, the target products include any one of the actual products; Perform similarity detection based on the product information corresponding to each group of target products to obtain a duplicate detection result; Based on the duplicate detection result corresponding to each cluster, perform a deduplication operation on the actual products to obtain deduplicated actual products.
2. The data processing method according to claim 1, characterized in that The calculating the similarity between different actual products based on the product information includes: Preprocess the product information to generate preset information; wherein, the preprocessing includes one or more of stop word removal operation, filtering operation, and text format conversion operation; Extract features from the preset information to obtain a feature vector corresponding to the preset information; Based on the feature vector corresponding to each actual product, determine the similarity between different actual products.
3. The data processing method according to claim 2, wherein The determining the similarity between different actual products based on the feature vector corresponding to each actual product includes: Based on the feature vector corresponding to each actual product, determine the cosine similarity between the feature vectors of different actual products; Use the cosine similarity as the similarity between the feature vectors of different actual products.
4. The data processing method according to claim 2, wherein The determining the similarity between different actual products based on the feature vector corresponding to each actual product includes: Based on the feature vector corresponding to each actual product, determine the target distance between the feature vectors of different actual products; wherein, the target distance includes any one of Euclidean distance and Manhattan distance; Use the target distance as the similarity between the feature vectors of different actual products.
5. The data processing method according to claim 1, wherein The determining at least one group of target products in each cluster based on the similarity includes: Use the actual products in each cluster with a similarity greater than the similarity threshold as a group of target products.
6. The data processing method according to claim 1, characterized in that, The performing similarity detection on the target products based on the product information of each group to obtain a duplicate detection result includes: Input the product information of each group of target products into a detection model for similarity detection to obtain a duplicate detection result; wherein, the detection model is trained based on the product information of historical similar products.
7. The data processing method according to claim 1, wherein The duplicate detection result includes: product duplication; The performing a deduplication operation on the actual products based on the duplicate detection result corresponding to each cluster to obtain deduplicated actual products includes: Based on the duplicate detection result corresponding to each cluster, delete any actual product with a duplicate detection result of product duplication to obtain deduplicated actual products.
8. An electronic device, characterized in that, Including: A memory and a processor, the memory being used for storing a computer program; the processor being used for, when executing the computer program, enabling the electronic device to implement the data processing method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by the processor, they are used to implement the data processing method according to any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product runs on a computer, it causes the computer to execute the data processing method according to any one of claims 1-7.