Multi-dimensional auto parts big data processing method and system based on vertical model
By employing a multi-dimensional auto parts big data processing method based on vertical class models, text feature vectors and semantic clustering are constructed, solving the accuracy and efficiency problems in auto parts data processing and achieving efficient deduplication and optimized data utilization of auto parts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG YUANDUN NETWORK TECH CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-05-29
AI Technical Summary
Existing auto parts data processing methods are insufficient to accurately depict parts compatibility relationships and regional price differences, resulting in limited data utilization depth. Furthermore, existing methods for identifying the same auto parts are not accurate enough, affecting the accuracy and reliability of deduplication.
A multi-dimensional auto parts big data processing method based on vertical class models is adopted. By constructing text feature vectors of auto parts, analyzing attribute similarity and importance, and combining semantic clustering, auto parts are classified and deduplicated. The vertical class model is used to evaluate the similarity of auto parts to improve the accuracy and efficiency of data processing.
It improves the accuracy of auto parts similarity analysis, reduces redundant data, lowers storage pressure, and enhances data utilization efficiency and deduplication accuracy.
Smart Images

Figure CN121919215B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a multi-dimensional auto parts big data processing method and system based on a vertical category model. Background Technology
[0002] With the continuous growth in the number of vehicles and parts, the scale of auto parts data is constantly expanding, and the data sources are diverse, the structure is complex, and the correlations are becoming increasingly strong. Existing auto parts data processing methods mostly rely on general models or rule analysis, which makes it difficult to accurately characterize parts compatibility relationships and regional price differences, and the depth of data utilization is limited, making it difficult to support real-time business decisions. Vertical models, by integrating industry knowledge and business rules, can improve the ability to understand the characteristics and inherent relationships of multi-dimensional auto parts data.
[0003] Due to the massive scale of auto parts data, redundant data for the same auto parts needs to be removed in practical applications to reduce storage pressure. However, existing technologies typically rely on simple similarity analysis to determine whether different auto parts data can be considered the same auto parts, without fully considering the differences in description of the same auto parts from different data sources. This can easily lead to different auto parts being misclassified as the same or the same auto parts being misclassified as different, thus affecting the accuracy and reliability of auto parts data deduplication. Summary of the Invention
[0004] To address the aforementioned technical problems, the purpose of this application is to provide a multi-dimensional auto parts big data processing method and system based on a vertical category model. The specific technical solution adopted is as follows:
[0005] In a first aspect, embodiments of this application provide a multi-dimensional auto parts big data processing method based on a vertical category model, the method comprising the following steps:
[0006] Periodically acquire multidimensional data information of auto parts; construct text feature vectors of auto parts based on key text information in the multidimensional data information;
[0007] From historically acquired auto parts, filter out auto parts with the same name for each auto part acquired in the current acquisition; analyze the similarity between each attribute of each auto part and each attribute of its corresponding auto parts, and determine the first similarity between each attribute of each auto part and its corresponding auto parts.
[0008] Construct a first similarity time series between each attribute of each auto part and all auto parts with the same name; determine the attribute importance of each attribute of each auto part by the degree of synchronous change between each attribute of each auto part and the time series corresponding to its remaining attributes, and by the numerical distribution of the time series.
[0009] The similarity of text feature vectors between any two auto parts in a single acquisition is evaluated based on a vertical classification model to classify the auto parts. The first similarity between each attribute of any auto part in the same category and the remaining auto parts is determined. Combined with the attribute importance of each attribute of the auto part, the auto part similarity between the auto part and the remaining auto parts is calculated to remove duplicates of the same category of auto parts.
[0010] In one embodiment, the key text information includes the vehicle brand, vehicle series, and model number of the auto parts.
[0011] In one embodiment, determining the first similarity between each attribute of the auto parts and an auto part with the same name includes:
[0012] Calculate the similarity between any attribute of each auto part and any attribute of auto parts with the same name, and take the maximum value of all similarity values corresponding to any attribute as the first similarity between any attribute and auto parts with the same name.
[0013] In one embodiment, determining the importance of the attribute includes:
[0014] Identify each mutation point in the first similarity time series, count the number of mutation points in the time series corresponding to any attribute of each auto part and its remaining attributes that have the same position, and determine the fusion result of the number of mutation points between any attribute of each auto part and its remaining attributes.
[0015] Based on the fusion results and the numerical distribution, the attribute importance of any attribute of each auto part is determined.
[0016] In one embodiment, the mean of the time series corresponding to any attribute of each auto part is calculated, and the attribute importance of any attribute of each auto part is positively correlated with the mean and the fusion result.
[0017] In one embodiment, the classification of auto parts includes:
[0018] By using a vertical classification model to convert text feature vectors into high-dimensional semantic vectors, clustering is performed on all the high-dimensional semantic vectors of auto parts acquired in a single transaction to obtain the classification results of all the auto parts acquired in a single transaction.
[0019] In one embodiment, the calculation process for the auto parts similarity is as follows:
[0020] The differences in the number of attributes between any given auto parts and the remaining auto parts in the same category are statistically analyzed; the first similarity between each attribute of the given auto parts and the remaining auto parts is denoted as the category similarity; and the product of the attribute importance of each attribute of the given auto parts and the category similarity corresponding to that attribute is calculated.
[0021] By combining the product of all attributes corresponding to any given auto part, taking into account the difference in the number of attributes, and the similarity of the given auto part with the remaining auto parts in terms of auto part name, the similarity between the given auto part and the remaining auto parts in the same category is calculated.
[0022] In one embodiment, the similarity of the auto parts is positively correlated with the product and the similarity of the name of any auto part to the remaining auto parts, and negatively correlated with the difference in the number of attributes.
[0023] In one embodiment, the deduplication process for similar auto parts includes:
[0024] Threshold segmentation is performed on the similarity between any two auto parts of the same type to subdivide the same type of auto parts into groups of auto parts; duplicate auto parts in the same group are removed based on the similarity between the attributes of the auto parts.
[0025] Secondly, embodiments of this application also provide a multi-dimensional auto parts big data processing system based on a vertical category model, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of any of the methods described above.
[0026] This application has at least the following beneficial effects:
[0027] This application constructs text feature vectors based on key textual information in auto parts, which helps to capture the core attributes and characteristics of auto parts, enhances the depth of data analysis, and makes subsequent similarity analysis between auto parts more accurate. By filtering historical data, it efficiently identifies auto parts with the same name for each currently acquired auto part, ensuring data consistency and reliability and avoiding errors in attribute importance calculation. Furthermore, by analyzing the degree of synchronous change between the time series of the first similarity corresponding to the attributes of each auto part, it effectively determines the importance of each attribute, improves the accuracy of attribute importance determination, and provides a reliable basis for assessing the similarity between different auto parts. This application improves data processing efficiency by classifying auto parts, while also facilitating accurate calculation of the similarity between auto parts. Finally, calculating the similarity between different auto parts can further narrow the identification range of duplicate auto parts, improve the accuracy of duplicate auto part identification, effectively deduplicate auto parts, reduce redundant data, improve data utilization efficiency, and reduce data storage pressure. Attached Figure Description
[0028] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 A flowchart illustrating the steps of a multi-dimensional auto parts big data processing method based on a vertical category model, provided in one embodiment of this application. Detailed Implementation
[0030] To further illustrate the technical means and effects adopted by this application to achieve the intended purpose of the invention, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the multi-dimensional auto parts big data processing method and system based on a vertical category model proposed in this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0032] The following, in conjunction with the accompanying drawings, details the specific scheme of the multi-dimensional auto parts big data processing method and system based on the vertical category model provided in this application.
[0033] Please see Figure 1 , which shows the step flow chart of the multi-dimensional auto parts big data processing method provided by an embodiment of the present application. The method includes the following steps:
[0034] S1. Regularly obtain the multi-dimensional data information of auto parts; construct the text feature vector of auto parts according to the key text information in the multi-dimensional data information.
[0035] Access the auto parts trading platforms of each auto parts supplier through the data interface, and obtain various types of data uploaded to the auto parts trading platform on the same day every day. For the data collected every day, according to the "Industrial Internet Identification Resolution - Automotive Parts - Identification Coding", the data containing the part identification coding is retained, and the data not containing the part identification coding is filtered, so as to limit the data analyzed in this embodiment to the multi-dimensional data information of auto parts. Then, attribute extraction is performed on the multi-dimensional data information of auto parts screened out every day through regular expressions. In this embodiment, the attributes of auto parts include the vehicle brand, vehicle series, model, price, color, size, and place of origin data of the auto parts. Implementers can set the attribute types of auto parts according to the actual situation, and this embodiment does not limit this. In addition, the multi-dimensional data information of auto parts contains the name of the auto parts.
[0036] It should be noted that: since the vehicle brands in different data sources may be abbreviations or English names, after the auto parts data is collected, all abbreviations or English names are first changed to Chinese names. For example, "BMW" is changed to the Chinese name "BMW" to facilitate the subsequent unified coding of auto parts attributes and avoid coding errors caused by different expressions of the same meaning attributes.
[0037] S2. Screen the auto parts with the same name as the auto parts obtained this time from the auto parts obtained historically; analyze the similarity between the attributes of each auto part and the attributes of its auto parts with the same name, and determine the first similarity between the attributes of each auto part and its auto parts with the same name.
[0038] Although there may be multiple naming methods for auto parts in different regions or different business scenarios, in the large-scale auto parts data, there will still objectively be a certain number of auto parts data with exactly the same name; based on this, taking the i-th auto part collected this time as an example in this embodiment, from the auto parts data collected in the previous time of this time, obtain T auto parts with exactly the same name as the i-th auto part in the most recent time period, and all of them are used as the auto parts with the same name of the i-th auto part. It should be noted that when the number of auto parts with the same name of the i-th auto part is less than T, then obtain all the auto parts with the same name of the i-th auto part. In this embodiment, T = 100 is set, and implementers can set it according to the actual situation, and this embodiment does not limit this.
[0039] The specific method for obtaining auto parts with the same name is as follows:
[0040] In this embodiment, the Jaccard similarity between the name of the i-th auto part and the historical name of the h-th auto part is calculated. If the value is 1, it means that the name of the i-th auto part is the same as that of the h-th auto part, and the h-th auto part is an auto part with the same name as the i-th auto part. Otherwise, the h-th auto part is not an auto part with the same name as the i-th auto part.
[0041] In another embodiment, a brute-force matching algorithm is used to match the names of the i-th auto part and the h-th auto part. If the match is successful, it is determined that the names of the i-th auto part and the h-th auto part are completely identical, and the h-th auto part is an auto part with the same name as the i-th auto part. Otherwise, the h-th auto part is not an auto part with the same name as the i-th auto part.
[0042] Considering that even for the same type of auto parts, some attributes may differ due to region or other factors, and that in the multidimensional data of auto parts, some attributes do not change independently but rather exhibit stable covariance relationships—that is, a change in one attribute may simultaneously cause changes in multiple attributes, such as changes in price, color, and place of origin caused by material—this embodiment evaluates the importance of attributes to auto parts by analyzing the degree of attribute variation.
[0043] However, considering that some attributes, besides the name, may have the same meaning but different names in different data sources, it is necessary to conduct in-depth analysis of each attribute in conjunction with the vertical category model. Take the u-th attribute of the i-th auto part and the p-th auto part with the same name as the i-th auto part as an example:
[0044] In this embodiment, the text data corresponding to the u-th attribute of the i-th auto part and the attributes of the p-th auto part with the same name are converted into word vector format using an auto parts vertical category model. Then, a similarity algorithm is used to calculate the similarity between the word vector corresponding to the u-th attribute and the word vectors corresponding to the attributes of the p-th auto part with the same name, and the maximum similarity is taken as the first similarity between the u-th attribute of the i-th auto part and the p-th auto part with the same name. The similarity algorithm is not limited to cosine similarity or edit similarity; this embodiment uses cosine similarity.
[0045] In another embodiment, the similarity between the u-th attribute of the i-th auto part and each attribute of the p-th auto part with the same name can be calculated sequentially using the Synonyms toolkit, and the maximum similarity is taken as the first similarity between the u-th attribute of the i-th auto part and the p-th auto part with the same name.
[0046] The first similarity is used to reflect the semantic correspondence of the u-th attribute of the i-th auto part in the attribute set of the p-th auto part with the same name, and can represent the identifiability and consistency of the attribute under different data sources.
[0047] S3, construct the first similarity time series between each attribute of each auto part and all auto parts with the same name; determine the attribute importance of each attribute of each auto part by the degree of synchronous change between each attribute of each auto part and the time series corresponding to its remaining attributes, and the numerical distribution of the time series.
[0048] The first similarity between the u-th attribute of the i-th auto part and all auto parts with the same name is arranged according to the collection time sequence of the auto parts with the same name, forming the time sequence sequence of the first similarity corresponding to the u-th attribute of the i-th auto part. Accordingly, the time sequence sequence of the first similarity corresponding to each attribute of each auto part can be obtained.
[0049] This embodiment identifies each mutation point in the first similarity time series using a mutation point detection algorithm and records the position of each mutation point in its respective first similarity time series. For the position of each mutation point in the first similarity time series corresponding to the u-th attribute of the i-th auto part, the number of mutation points in the first similarity time series corresponding to the remaining attributes of the i-th auto part that share the same position as the mutation points in the first similarity time series corresponding to the u-th attribute is counted. The fusion result of the number of mutation points corresponding to all other attributes of the i-th auto part is determined, and the mean of all elements in the first similarity time series corresponding to the u-th attribute is calculated and denoted as the first mean. A larger first mean indicates a more stable semantic expression of the u-th attribute among auto parts with the same name, less affected by regional or naming differences, and greater importance in subsequent auto part matching. Based on the first mean and the fusion result, the attribute importance of the u-th attribute of the i-th auto part is determined.
[0050] It should be noted that fusion refers to combining multiple variables, which can be calculated using methods such as addition, multiplication, or averaging. In this embodiment, the fusion result is calculated using the averaging method. The fusion result reflects the degree to which other attributes change synchronously at the same time when the u-th attribute of the i-th auto part changes; the larger the value, the more attributes change due to the linkage effect when the u-th attribute changes, and thus the greater the importance of the u-th attribute to the i-th auto part.
[0051] Furthermore, the fusion result corresponding to all attributes of the i-th auto part is normalized using a maximum-minimum normalization method to avoid the fusion result being too large and weakening the influence of other parameters.
[0052] In this embodiment, the sum of the normalized value of the fusion result corresponding to the u-th attribute of the i-th auto part and the first mean of the u-th attribute is used as the attribute importance of the u-th attribute of the i-th auto part. The greater the attribute importance, the greater the semantic correspondence of the u-th attribute among all auto parts with the same name, and the stronger the linkage effect on other attributes when it changes, thus playing a more crucial role in subsequently measuring the matching degree between auto parts.
[0053] S4. Based on the vertical classification model, evaluate the similarity of text feature vectors between any two auto parts in a single acquisition and classify the auto parts; determine the first similarity between each attribute of any auto part in the same category and the remaining auto parts, and calculate the auto part similarity between any auto part and the remaining auto parts by combining the attribute importance of each attribute of the auto part, so as to remove duplicates of the same category of auto parts.
[0054] Because auto parts belong to the automotive aftermarket, their naming conventions vary significantly across different regions, companies, and supply chain systems due to historical development and industry structure. Therefore, analyzing auto parts solely based on their names can easily lead to significant biases, affecting the reliability of subsequent analysis results and the effectiveness of business decisions. Thus, it is necessary to combine vertical models within the auto parts sector for more in-depth analysis and processing of auto parts data.
[0055] Given that there is usually a clear and stable compatibility relationship between auto parts and specific car models, the parts of different car models often have essential differences in terms of structural dimensions, installation position and function. If all auto parts are analyzed in a unified manner, it will easily lead to a large number of invalid comparisons. Therefore, this embodiment classifies and processes all auto parts data acquired on the same day.
[0056] Taking the a-th auto part as an example, based on the vehicle brand, series, and model data of the a-th auto part, a text feature vector is constructed for the a-th auto part, specifically: [vehicle brand, series, model]. Similarly, text feature vectors are constructed for each auto part in the collected data.
[0057] A Transform-based automotive parts vertical model is used to sequentially convert the textual feature vectors of each automotive part into high-dimensional semantic vectors. Then, a semantic clustering method is used to cluster the high-dimensional semantic vectors of all automotive parts for that day, thereby achieving differentiation between different types of automotive parts. In this embodiment, the semantic clustering method is the k-means algorithm, and the value of k is determined using the elbow method. Implementers can choose other algorithms. In other embodiments, the automotive parts vertical model can also be a vertical BERT model, a vertical RoBERTa model, etc.
[0058] After analyzing the importance of attributes to auto parts, duplicate auto parts in the same category obtained after clustering can be removed, thereby reducing data storage pressure. Taking the i-th and j-th auto parts in the same category as an example, since the auto part name is usually a direct description of the auto part type and function, the consistency of the auto part name can be analyzed first.
[0059] In this embodiment, the names of the i-th and j-th auto parts are compared using the method of obtaining auto parts with the same name. If the names of the two auto parts are completely identical, the i-th and j-th auto parts are directly classified into the same group of auto parts.
[0060] If the names of the i-th auto parts and the j-th auto parts are not completely identical, then the similarity between the names of the i-th auto parts and the j-th auto parts is first calculated. Specifically, the word vector of the name of the i-th auto parts and the word vector of the name of the j-th auto parts are calculated to reflect the semantic closeness of the two auto parts names. In this embodiment, cosine similarity is used for calculation. Implementers can choose other existing feasible similarity calculation methods.
[0061] Furthermore, considering that the names of auto parts may differ due to different business scenarios or data sources, but for the same auto part, the attribute information it contains is usually highly consistent, this embodiment continues to analyze the attribute characteristics of auto parts.
[0062] Since the number of attributes contained in the same type of auto parts usually remains relatively stable, this embodiment calculates the difference in the number of attributes between the i-th auto part and the j-th auto part to reflect the degree of difference between the two auto parts in terms of attribute structure and scale. Here, the difference represents the degree of difference between two variables, which can be calculated using methods such as the absolute value of the difference or the square of the difference. In this embodiment, the difference in the number of attributes is the absolute value of the difference in the number of attributes.
[0063] Furthermore, for the u-th attribute of the i-th auto part, this embodiment uses the same calculation method as the first similarity calculation between the u-th attribute and auto parts with the same name to calculate the first similarity between the u-th attribute and the j-th auto part, denoted as the class similarity. The product of the attribute importance of the u-th attribute of the i-th auto part and the class similarity is calculated, and this product is used as the attribute correspondence between the u-th attribute of the i-th auto part and the j-th auto part. The attribute correspondence reflects the semantic matching strength of the two auto parts on the u-th attribute and its importance weight in the overall attribute structure.
[0064] Similarly, the attribute correspondence between each attribute of the i-th auto part and the j-th auto part is calculated sequentially. The average of these attribute correspondences for all attributes of the i-th auto part is then used as the second average between the i-th and j-th auto parts. This second average reflects the overall matching degree of the two auto parts across all attribute dimensions. A larger second average indicates greater consistency in the overall structural attributes of the two auto parts, and a higher probability that they are duplicate auto parts.
[0065] Based on the above analysis, this embodiment calculates the auto parts similarity between the i-th auto parts and the j-th auto parts, specifically expressed as: ;
[0066] In the formula, This represents the similarity between the i-th and j-th auto parts in the same category. This represents the degree of similarity in name between the i-th and j-th auto parts in the same category. This is the second mean between the i-th and j-th auto parts in the same category. This represents the difference in the number of attributes between the i-th and j-th auto parts of the same type. These are preset parameter tuning coefficients that are greater than 0 and are integers, used to avoid denominators of 0. In this embodiment... The implementer can set it according to the actual situation, and this embodiment does not impose any restrictions on it.
[0067] Auto parts similarity reflects the degree of overall consistency between the i-th auto parts and the j-th auto parts at the semantic and attribute structure levels. The greater the auto parts similarity, the higher the degree of matching between the two auto parts in terms of name semantics, functional attributes, and key attribute structure, and thus the greater the possibility that they are duplicate auto parts.
[0068] Furthermore, in all the data collected in this instance, the similarity between the i-th auto part and all other auto parts of the same type is calculated. The similarity between the i-th auto part and all other auto parts is sorted in descending order. The sorted auto part similarity data is segmented using the natural breakpoint method. The auto parts corresponding to the auto parts with the largest average similarity after segmentation are recorded as the same group of auto parts as the i-th auto part in this data collection.
[0069] For other auto parts of the same type that have not yet been identified in the same group, perform the identification process for the same group of auto parts as described above until all auto parts are grouped and distinguished.
[0070] It should be noted that once a certain auto part has been identified as a part in the same group, that auto part will no longer participate in the subsequent group identification process, thereby ensuring that each auto part is identified in the same group only once, so as to improve the overall processing efficiency.
[0071] For the i-th and m-th auto parts in the same group, if the first similarity between all attributes of the m-th auto part and the i-th auto part is greater than a preset threshold, it means that all attributes of the m-th auto part have a corresponding counterpart in the i-th auto part. In this case, the m-th auto part is considered a duplicate of the i-th auto part and is removed, thereby reducing the storage pressure on the system. The remaining auto parts in the same group are processed in the same way to achieve unified deduplication of auto parts data. A higher preset threshold indicates higher accuracy in identifying duplicate auto parts, but may result in missed identifications. Conversely, a lower preset threshold may lead to false identifications of duplicate auto parts. Experimentally, this embodiment sets the preset threshold to 0.8. Implementers can set it according to their actual situation; this embodiment does not impose any restrictions. It should be noted that for the identification of auto parts in the same group within the same type of auto parts, and for the identification of duplicate auto parts within the same group, this embodiment operates according to the order in which the auto parts were collected.
[0072] Based on the same inventive concept as the above methods, this application also provides a multi-dimensional auto parts big data processing system based on a vertical category model, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-described multi-dimensional auto parts big data processing methods based on a vertical category model.
[0073] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments of this specification have been described above. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0074] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0075] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A multidimensional auto parts big data processing method based on a vertical category model, characterized in that, The method includes the following steps: Periodically acquire multidimensional data information of auto parts; construct text feature vectors of auto parts based on key text information in the multidimensional data information; From historically acquired auto parts, filter out auto parts with the same name for each auto part acquired in the current acquisition; analyze the similarity between each attribute of each auto part and each attribute of its corresponding auto parts, and determine the first similarity between each attribute of each auto part and its corresponding auto parts. Construct a first similarity time series between each attribute of each auto part and all auto parts with the same name; determine the attribute importance of each attribute of each auto part by the degree of synchronous change between each attribute of each auto part and the time series corresponding to its remaining attributes, and by the numerical distribution of the time series. The similarity of text feature vectors between any two auto parts in a single acquisition is evaluated based on the vertical classification model to classify the auto parts; the first similarity between each attribute of any auto part in the same category and the remaining auto parts is determined; and the auto part similarity between any auto part and the remaining auto parts is calculated by combining the attribute importance of each attribute of the auto part, so as to remove duplicates of the same category of auto parts. Determining the importance of the attribute includes: Identify each mutation point in the first similarity time series, count the number of mutation points in the time series corresponding to any attribute of each auto part and its remaining attributes that have the same position, and determine the fusion result of the number of mutation points between any attribute of each auto part and its remaining attributes. Based on the fusion results and the numerical distribution, the attribute importance of any attribute of each auto part is determined; The calculation process for the similarity of the auto parts is as follows: The differences in the number of attributes between any given auto parts and the remaining auto parts in the same category are statistically analyzed; the first similarity between each attribute of the given auto parts and the remaining auto parts is denoted as the category similarity; and the product of the attribute importance of each attribute of the given auto parts and the category similarity corresponding to that attribute is calculated. By combining the product of all attributes corresponding to any given auto part, taking into account the difference in the number of attributes, and the similarity of the given auto part with the remaining auto parts in terms of auto part name, the similarity between the given auto part and the remaining auto parts in the same category is calculated.
2. The multi-dimensional auto parts big data processing method based on a vertical category model as described in claim 1, characterized in that, The key text information includes the vehicle brand, model, and vehicle type of the auto parts.
3. The multidimensional auto parts big data processing method based on a vertical category model as described in claim 1, characterized in that, The determination of the first similarity between each attribute of each auto part and its corresponding auto part includes: Calculate the similarity between any attribute of each auto part and any attribute of auto parts with the same name, and take the maximum value of all similarity values corresponding to any attribute as the first similarity between any attribute and auto parts with the same name.
4. The multi-dimensional auto parts big data processing method based on a vertical category model as described in claim 1, characterized in that, Calculate the mean of the time series corresponding to any attribute of each auto part. The attribute importance of any attribute of each auto part is positively correlated with the mean and the fusion result.
5. The multidimensional auto parts big data processing method based on a vertical category model as described in claim 1, characterized in that, The classification of auto parts includes: By using a vertical classification model to convert text feature vectors into high-dimensional semantic vectors, clustering is performed on all the high-dimensional semantic vectors of auto parts acquired in a single transaction to obtain the classification results of all the auto parts acquired in a single transaction.
6. The multidimensional auto parts big data processing method based on a vertical category model as described in claim 1, characterized in that, The similarity of the auto parts is positively correlated with the product and the similarity of the name of any auto part to the remaining auto parts, and negatively correlated with the difference in the number of attributes.
7. The multidimensional auto parts big data processing method based on a vertical category model as described in claim 1, characterized in that, The deduplication process for similar auto parts includes: Threshold segmentation is performed on the similarity between any two auto parts of the same type to subdivide the same type of auto parts into groups of auto parts; duplicate auto parts in the same group are removed based on the similarity between the attributes of the auto parts.
8. A multi-dimensional auto parts big data processing system based on a vertical category model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Coal machine spare part similarity identification and coding unification system and method and storage medium
CN121434738A