Product clustering analysis method based on deep learning model

By constructing a named pattern recognition model and performing semantic analysis, the problem of low accuracy of classification analysis of product names in manufacturing industry is solved, and a higher accuracy of classification analysis is achieved.

CN118820813BActive Publication Date: 2025-05-16CHINA NAT INST OF STANDARDIZATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410894273.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-04
Publication Date
2025-05-16
Estimated Expiration
2044-07-04

AI Technical Summary

Technical Problem

In the prior art, the name classification analysis of manufacturing products has low accuracy and it is difficult to accurately identify product name categories.

Method used

By constructing a naming pattern recognition model, the naming pattern of the product name to be analyzed is obtained, and converted into a numerical vector, semantic analysis and cluster analysis are performed, the product name feature vector and similar feature index are obtained, and quality evaluation and optimization are finally carried out to obtain the first cluster analysis results.

Benefits of technology

It improves the accuracy of name classification analysis of manufacturing products and effectively solves the problem of low name classification analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118820813B_ABST
    Figure CN118820813B_ABST
Patent Text Reader

Abstract

The present invention discloses a product clustering analysis method based on a deep learning model, and relates to the technical field of electrical digital data processing. The product clustering analysis method based on a deep learning model comprises the following steps: model construction; numerical vector conversion; clustering analysis result acquisition; quality assessment. The present invention acquires the naming pattern of the product name to be analyzed by constructing a naming pattern recognition model, and converts the product name to be analyzed acquired in real time into a numerical vector of the name to be analyzed, and then obtains a similarity feature index based on the extracted product name feature vector, and then obtains the clustering analysis result based on the product name feature vector, and finally obtains a quality assessment index for the clustering analysis result, and obtains the first clustering analysis result based on the quality assessment index, thereby improving the accuracy of the name classification analysis of manufacturing products and solving the problem of low accuracy of the name classification analysis of manufacturing products in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a product clustering analysis method based on a deep learning model. Background Art

[0002] With the development of the Internet, product clustering analysis has shown unprecedented value in the manufacturing industry. The manufacturing product clustering analysis method not only uses big data and machine learning technology, but also integrates the real-time and massive information of the Internet to achieve accurate capture and classification of product attributes. These methods can conduct in-depth mining and clustering based on multivariate data such as product design, function, and user feedback, helping companies quickly identify similar product groups in the market, understand consumer demand, and optimize product portfolios and supply chain strategies. At the same time, clustering analysis also provides a more accurate market positioning for the manufacturing industry, helping companies stand out in the fierce competition.

[0003] Existing name classification analysis methods are mainly based on data mining and machine learning technology, which can effectively classify products by calculating the similarity between product data. These methods include K-means clustering, hierarchical clustering, DBSCAN, etc., which can automatically classify similar products into one category based on multiple product attributes, features, or user evaluation dimensions. These clustering methods can reveal the inherent relationship between products and help companies identify the similarities and differences between products, thereby optimizing product positioning, formulating marketing strategies, and performing market segmentation. Through cluster analysis, companies can better understand market demand and competitive trends, and improve market competitiveness and business efficiency.

[0004] For example, the invention patent with announcement number: CN103617163B discloses a method for rapid target association based on cluster analysis, comprising: first, constructing a structure array in a rapid association module, the number of array elements of which is equal to the sum of the number of targets detected by two sub-sources, each element in the structure array represents a target detected by a certain sub-source, setting the initial association state of all targets to be unassociated, and marking in the structure array which sub-source the target is measured from; secondly, recursively performing rapid sorting and rapid clustering on the targets to be associated in the above structure array according to the established priority hierarchy, using the position components in the Earth-centered Earth-fixed ECEF coordinate system: X, Y and Z coordinate values ​​and attribute information as keywords, respectively, to construct a hierarchically ordered spatial index tree; then, with the support of the spatial index tree, using a density-based clustering algorithm to recursively cluster the targets with the same attributes in adjacent positions of the two sub-sources, and when the total number of sub-sequence targets is less than a preset value k or there is no unused position or attribute information, a one-to-one comparison association method is used to complete the final association judgment of the target.

[0005] For example, the parallelization of large-scale data clustering analysis announced in the invention patent with announcement number: CN102855259B includes: a cluster selector, which is configured to determine multiple sample clusters, and reproduce the multiple sample clusters at each of multiple processing cores; a sample divider, which is configured to divide multiple samples with associated attributes stored in a database into sample subsets whose number corresponds to the number of multiple processing cores, and is also configured to associate each of the number of sample subsets with a corresponding one of the multiple processing cores; an integration operator, which is configured to perform a comparison of each sample relative to each of the multiple sample clusters reproduced at the corresponding processing core based on the associated attributes of each sample in each sample subset at each corresponding core of the multiple processing cores; and an attribute divider, which is configured to divide the attributes associated with each sample into attribute subsets so that they can be processed in parallel during the comparison.

[0006] However, in the process of implementing the technical solution of the invention in the embodiments of the present application, the present application found that the above technology has at least the following technical problems:

[0007] In the prior art, traditional product name classification and analysis methods are based on text matching and keyword search. When processing manufacturing products, due to the diversity and professionalism of product names, it is often difficult to accurately identify the product name category, and there is a problem of low accuracy in the name classification and analysis of manufacturing products. Summary of the invention

[0008] The embodiment of the present application solves the problem of low accuracy of name classification analysis of manufacturing products in the prior art by providing a product clustering analysis method based on a deep learning model, thereby achieving the effect of improving the accuracy of name classification analysis of manufacturing products.

[0009] The embodiment of the present application provides a product clustering analysis method based on a deep learning model, comprising the following steps: S1, collecting product names of manufacturing products from a product name database and constructing a naming pattern recognition model, wherein the naming pattern recognition model is used to recognize the naming pattern of the product names to be analyzed acquired in real time; S2, acquiring the naming pattern of the product names to be analyzed in combination with the naming pattern recognition model, and converting the product names to be analyzed acquired in real time into numerical vectors of the names to be analyzed; S3, performing semantic analysis on the numerical vectors of the names to be analyzed and extracting the product name feature vectors of the numerical vectors of the names to be analyzed in combination with the clustering target to obtain a similarity feature index, and obtaining a clustering analysis result based on the product name feature vectors, wherein the similarity feature index is used to describe the degree of similarity between the product names to be analyzed and between the product clusters to be analyzed; S4, performing quality assessment on the clustering analysis results to obtain a quality assessment index, optimizing the clustering algorithm based on the quality assessment index, and then repeatedly executing S1 until the quality assessment index meets the corresponding threshold and obtains the first clustering analysis result, wherein the quality assessment index is used to describe the quality of the clustering analysis of the product name feature vectors.

[0010] Furthermore, the specific steps of identifying the naming pattern are: using an entity test set to perform a model test on the training model to obtain an entity test index, and the entity test index is used to test the performance of the training model; judging whether the obtained entity test index meets the test threshold, if so, identifying the named entity through the training model, if not, optimizing the parameters of the training model; outputting the identified named entity and the corresponding named entity type to obtain the identified naming pattern.

[0011] Furthermore, the method for obtaining the entity test index is as follows: extracting an entity test score from the model testing process, the entity test score including the learning rate, the accuracy rate and the recall rate, and obtaining the corresponding preset standard learning rate according to the learning rate; obtaining the entity test index of the training model according to the obtained entity test score and the preset standard learning rate; the entity test index is calculated using the following formula:

[0012]

[0013] Where e is a natural constant, g is the number of the training round, g = 1, 2, ..., G, G is the total number of training rounds, ST represents the entity test index of the training model, X g represents the learning rate of the training model in the gth training round, α g represents the accuracy of the training model in the gth training round, β g represents the recall rate of the training model in the gth training round, and X0 represents the preset standard learning rate of the training model.

[0014] Furthermore, the method for obtaining the similarity feature index is as follows: obtaining the similarity feature scores between the product names to be analyzed based on the extracted product name feature vectors, the similarity feature scores including semantic similarity and string similarity; obtaining the similarity feature score vector based on the obtained similarity feature scores, and simultaneously obtaining its corresponding reference vector, the similarity feature score vector including a semantic similarity vector and a string similarity vector; obtaining the similarity feature index of the product to be analyzed based on the obtained similarity feature score vector and its corresponding reference vector.

[0015] Furthermore, the similarity feature index is calculated using the following formula:

[0016]

[0017] In the formula, m represents the number of the product name to be analyzed, m = 1, 2, ..., M, M is the total number of product names to be analyzed, XS m represents the similarity feature index of the mth product name to be analyzed, represents the semantic similarity vector of the mth product name to be analyzed, represents the reference semantic similarity vector of the mth product name to be analyzed, represents the string similarity vector of the mth product name to be analyzed, Represents the reference string similarity vector of the mth product name to be analyzed.

[0018] Furthermore, the method for obtaining the quality assessment index is as follows: obtaining a cluster quality score based on the cluster analysis results, the cluster quality score including intra-cluster similarity and inter-cluster offset; obtaining a central cluster micro-number based on the obtained cluster quality score, the central cluster micro-number including a central similarity micro-number and a central offset micro-number; obtaining a quality assessment index of the cluster analysis based on the obtained cluster quality score and central cluster micro-number.

[0019] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0020] 1. The naming pattern of the product name to be analyzed is obtained by constructing a naming pattern recognition model, and the product name to be analyzed obtained in real time is converted into a numerical vector of the name to be analyzed, and then the similarity feature index is obtained based on the extracted product name feature vector, and then the cluster analysis result is obtained based on the product name feature vector, and finally the quality evaluation index is obtained for the cluster analysis result, and the first cluster analysis result is obtained based on the quality evaluation index, thereby realizing the cluster analysis of the names of manufacturing products, thereby improving the accuracy of the classification analysis of the names of manufacturing products, and effectively solving the problem of low accuracy of the classification analysis of the names of manufacturing products in the prior art.

[0021] 2. By extracting the entity test score from the model testing process, and obtaining the corresponding preset standard learning rate according to the learning rate, and then obtaining the entity test index of the training model according to the obtained entity test score and the preset standard learning rate, and finally testing the performance of the training model according to the obtained entity test index, the performance of the training model is quantified, thereby achieving a more accurate evaluation of the performance of the training model.

[0022] 3. Obtain similarity feature scores between product names to be analyzed based on the extracted product name feature vectors, obtain similarity feature score vectors based on the obtained similarity feature scores, and simultaneously obtain the corresponding reference vectors, and then obtain similarity feature indexes of products to be analyzed based on the obtained similarity feature score vectors and the corresponding reference vectors, and finally evaluate the similarity between product names to be analyzed and between product clusters to be analyzed based on the obtained similarity feature indexes, thereby realizing the quantification of the similarity between product names to be analyzed and between product clusters to be analyzed, and further realizing a more accurate evaluation of the similarity between product names to be analyzed and between product clusters to be analyzed. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A flow chart of a product clustering analysis method based on a deep learning model provided in an embodiment of the present application;

[0024] Figure 2 A flowchart for constructing a naming pattern recognition model provided in an embodiment of the present application;

[0025] Figure 3 A statistical graph of changes in the similarity feature index provided in the embodiments of the present application. DETAILED DESCRIPTION

[0026] The embodiment of the present application solves the problem of low accuracy of name classification analysis of manufacturing products in the prior art by providing a product clustering analysis method based on a deep learning model. The method collects product names of manufacturing products from a product name database and constructs a naming pattern recognition model. The naming pattern of the product name to be analyzed is then obtained in combination with the naming pattern recognition model. The product name to be analyzed obtained in real time is converted into a numerical vector of the name to be analyzed. Then, a semantic analysis is performed on the numerical vector of the name to be analyzed and a similarity feature index is obtained by combining the product name feature vector of the numerical vector of the name to be analyzed with the clustering target. At the same time, a clustering analysis result is obtained based on the product name feature vector. Finally, a quality assessment is performed on the clustering analysis result to obtain a quality assessment index. After the clustering algorithm is optimized based on the quality assessment index, the naming pattern recognition model is repeatedly constructed until the quality assessment index meets the corresponding threshold and the first cluster analysis result is obtained. This achieves the effect of improving the accuracy of name classification analysis of manufacturing products.

[0027] The technical solution in the embodiment of the present application is to solve the problem of low accuracy in the classification analysis of the names of the above-mentioned manufacturing products. The overall idea is as follows:

[0028] By constructing a naming pattern recognition model, the naming pattern of the product name to be analyzed is obtained, and the product name to be analyzed obtained in real time is converted into a numerical vector of the name to be analyzed. Then, the similarity feature index is obtained based on the extracted product name feature vector, and then the cluster analysis result is obtained based on the product name feature vector. Finally, the quality assessment index is obtained for the cluster analysis result, and the first cluster analysis result is obtained based on the quality assessment index, which improves the accuracy of name classification analysis of manufacturing products.

[0029] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0030] like Figure 1 As shown, it is a flow chart of a product clustering analysis method based on a deep learning model provided by an embodiment of the present application, and the method comprises the following steps: S1, collecting product names of manufacturing products from a product name database and constructing a naming pattern recognition model, the naming pattern recognition model is used to recognize the naming pattern of the product names to be analyzed acquired in real time; S2, obtaining the naming pattern of the product names to be analyzed in combination with the naming pattern recognition model, and converting the product names to be analyzed acquired in real time into a numerical vector of the names to be analyzed, the numerical vector of the product names to be analyzed is a numerical form that can be processed by a computer; S3, performing semantic analysis on the numerical vector of the names to be analyzed and combining it with the clustering target to obtain a naming pattern; The product name feature vector of the name numerical vector to be analyzed is obtained to obtain a similarity feature index, and at the same time, a cluster analysis result is obtained based on the product name feature vector. The similarity feature index is used to describe the similarity between the product names to be analyzed and between the product clusters to be analyzed. The product clusters to be analyzed represent a set of product names to be analyzed whose similarity reaches a corresponding threshold. S4, quality assessment is performed on the cluster analysis result to obtain a quality assessment index. After optimizing the clustering algorithm according to the quality assessment index, S1 is repeatedly executed until the quality assessment index meets the corresponding threshold and a first cluster analysis result is obtained. The quality assessment index is used to describe the quality of the cluster analysis of the product name feature vector.

[0031] In this embodiment, the product names of all manufacturing products are derived from the product catalogs and e-commerce websites of manufacturing companies. If a product name of a manufacturing product that does not exist in the product name database appears, it is stored in the product name database for product name update; the naming pattern represents the named entity and the corresponding named entity type, the product name database is used to collect the product names of all manufacturing products, the product name is a method of expressing the manufacturing product, and the named entity represents the essential information of the manufacturing product; the essential information of the manufacturing product generally includes the brand name, model and material, etc.; cluster analysis is used to classify the names of the products to be analyzed whose similar feature index meets the preset similarity threshold into the same category, and the cluster analysis result represents the category of the product name to be analyzed after the cluster analysis; the first cluster analysis result represents the category of the product name to be analyzed after optimization; the accuracy of the name classification analysis of the manufacturing products is improved.

[0032] Further, such as Figure 2 As shown, it is a flowchart for constructing a naming pattern recognition model provided in an embodiment of the present application, and the steps for constructing the naming pattern recognition model are: S11, dividing the product names after the named entities are annotated into an entity training set, an entity verification set and an entity test set according to a preset ratio, and the recurrent neural network model is a model with the function of recognizing entities; S12, setting the hyperparameters and training parameters of the selected recurrent neural network model, and using the entity training set to train the recurrent neural network model until the model converges to obtain a training model; S13, using the entity verification set to perform performance verification on the training model to obtain a performance verification result, and judging whether the performance verification result meets the corresponding threshold condition, if so, generating a naming pattern recognition model and identifying the naming pattern, otherwise retraining the model until the performance verification result meets the corresponding threshold condition.

[0033] In this embodiment, before dividing the product names according to a preset ratio, the collected product names of manufacturing products need to be cleaned and duplicate, erroneous and irrelevant information removed; when selecting a model, factors such as model performance, training time and model size need to be considered, and a recurrent neural network (RNN) is a type of neural network model mainly used to process sequence data, which is widely used in the field of natural language processing. Since the target requirement for the product names of manufacturing products is to process natural language, a recurrent neural network is selected here; hyperparameters are data for optimizing model performance, and training parameters are set data during the training process. Hyperparameters include the number of network layers and the number of hidden units, and training parameters include training rounds and batch size; performance verification results are used to verify the generalization ability of the training model, and the corresponding threshold conditions are set by preset personnel; the efficiency of building a naming pattern recognition model is improved, and the accuracy of name classification analysis of manufacturing products is further improved.

[0034] Furthermore, the specific steps for identifying the naming pattern are: use the entity test set to perform model testing on the training model to obtain the entity test index, and the entity test index is used to test the performance of the training model; determine whether the obtained entity test index meets the test threshold, if so, identify the naming pattern through the training model, if not, optimize the parameters of the training model; output the identified named entities and the corresponding named entity types to obtain the identified naming pattern.

[0035] In this embodiment, parameter optimization is performed by a genetic algorithm: first, an initial individual population is randomly generated, each individual represents a parameter combination, and a fitness evaluation is performed on each parameter combination to evaluate its performance in solving the problem. Selection is performed based on the fitness of the individual. The selection method commonly used is roulette selection or bidding selection. Excellent individuals have a higher probability of being selected to produce the next generation. Two or more individuals are selected from the selected individuals as parents, and offspring are generated through a crossover operation. The crossover process simulates the combination of genes in the biological world. The newly generated individuals are mutated, that is, random changes are introduced in the parameters of the individuals. The mutation operation helps to maintain the diversity of the population and avoid falling into a local optimal solution. The newly generated offspring replaces a part of the individuals in the original population. The replacement principle is to retain individuals with higher fitness. The above steps are repeated until the maximum number of iterations is met or the preset optimal solution is reached; when the algorithm terminates, the individual with the best fitness is returned as the optimal parameter combination; the accuracy of naming pattern recognition is improved, and the accuracy of name classification analysis of manufacturing products is further improved.

[0036] Furthermore, the method for obtaining the entity test index is as follows: extracting the entity test score from the model testing process, the entity test score includes the learning rate, the accuracy rate and the recall rate, and obtaining the corresponding preset standard learning rate according to the learning rate; obtaining the entity test index of the training model according to the obtained entity test score and the preset standard learning rate; the entity test index is calculated using the following formula:

[0037]

[0038] Where e is a natural constant, g is the number of the training round, g = 1, 2, ..., G, G is the total number of training rounds, ST represents the entity test index of the training model, X g represents the learning rate of the training model in the gth training round, α g represents the accuracy of the training model in the gth training round, β g represents the recall rate of the training model in the gth training round, and X0 represents the preset standard learning rate of the training model.

[0039] In this embodiment, the learning rate is the ratio of the data actually learned by the training model to all the training data, the accuracy is the ratio of the number of samples correctly predicted by the model to the total number of samples, and the sklearn.metrics.accuracy_score function can be used in Python to obtain the accuracy rate, the recall rate is the ratio of samples predicted by the model to the positive class among the samples that are actually in the positive class, and the sklearn.metrics.recall_score function can be used in Python to obtain the recall rate. The preset standard learning rate is set by the preset personnel, which is generally the average value of the learning rate in the actual situation; assuming G is 1, the statistical table of the change of the entity test index is shown in Table 1:

[0040] Table 1 Statistics of changes in entity test index

[0041]

[0042]

[0043] It can be seen from the above table that the entity test index cannot be directly obtained through a single data, but needs to be judged by the learning rate, and the accuracy and recall rate are comprehensively analyzed to obtain it. For example, when the learning rate of the first group of data is less than the preset standard learning rate, the entity test index is directly zero, and the accuracy of the second group is greater than that of the third group. However, the final entity test index of the second group is equal to the entity test index of the third group. Because the recall rate also has an impact, a comprehensive analysis is required; the performance of the test training model is quantified, which further improves the accuracy of the name classification analysis of manufacturing products.

[0044] Furthermore, the process of converting the product names to be analyzed obtained in real time into the numerical vectors of the names to be analyzed is as follows: deleting redundant symbols in the product names to be analyzed obtained in real time, and converting all the product names to be analyzed into lowercase, and the redundant symbols include special characters and punctuation marks; using a word segmentation tool to split the product names to be analyzed into separate words, and using a stem extraction method to restore the separate words until they become the root form of the words; obtaining a word prototype set and converting it into a bag of words, the word prototype set is a collection of prototypes of separate words, and the prototypes of separate words in the bag of words are the vectorized features of the product names to be analyzed; combining the bag of words model to map the product names to be analyzed with the numerical vector and normalizing the generated numerical vector to obtain the numerical vector of the name to be analyzed, and the bag of words model is used to convert the text of the product name to be analyzed into a numerical representation.

[0045] In this embodiment, redundant symbols in the product names to be analyzed obtained in real time are deleted until only the text content is retained, and all the product names to be analyzed are converted to lowercase to avoid the same word being regarded as different words due to different cases; the word segmentation tool uses space segmentation, for example, "Apple iPhone" can be segmented into ["apple", "iphone"]; the bag-of-words model represents each product name to be analyzed as a numerical vector, each element of the numerical vector represents the prototype of a single word in the bag-of-words, and its value is the frequency of occurrence or binary indication of the prototype of the single word in the product name to be analyzed; the generated numerical vectors are standardized to ensure that they have similar numerical ranges; through the above operations, accurate conversion of the numerical vectors of the names to be analyzed is achieved, and the accuracy of name classification analysis of manufacturing products is further improved.

[0046] Furthermore, a method for obtaining a similar feature index is as follows: obtaining similar feature scores between product names to be analyzed based on the extracted product name feature vector, wherein the similar feature scores include semantic similarity and string similarity; obtaining a similar feature score vector based on the obtained similar feature scores, and simultaneously obtaining its corresponding reference vector, wherein the similar feature score vector includes a semantic similarity vector and a string similarity vector; obtaining a similar feature index of the product to be analyzed based on the obtained similar feature score vector and its corresponding reference vector.

[0047] In this embodiment, the similarity between the names of products to be analyzed and between the clusters of products to be analyzed is quantified, which further improves the accuracy of the name classification analysis of manufacturing products.

[0048] Furthermore, the similarity feature index is calculated using the following formula:

[0049]

[0050] In the formula, m represents the number of the product name to be analyzed, m = 1, 2, ..., M, M is the total number of product names to be analyzed, XS m represents the similarity feature index of the mth product name to be analyzed, represents the semantic similarity vector of the mth product name to be analyzed, represents the reference semantic similarity vector of the mth product name to be analyzed, represents the string similarity vector of the mth product name to be analyzed, Represents the reference string similarity vector of the mth product name to be analyzed.

[0051] In this embodiment, the reference semantic similarity vector and the reference string similarity vector are both set by a preset person; the calculation of the similarity feature index is based on the dot product of the vectors, where and Represents the dot product of vectors, which is also called the scalar product. The result is the length of the projection of a vector in the direction of another vector, which is a scalar; defines the semantic similarity coefficient String similarity coefficient Then the similarity feature index can be simplified as: like Figure 3 As shown, it is a statistical graph of the changes in the similarity feature index provided in the embodiment of the present application. It can be seen from the graph that the semantic similarity coefficient and the string similarity coefficient are positively correlated with the similarity feature index. When the semantic similarity coefficient increases, the similarity feature index increases, and the similarity between the names of the products to be analyzed and between the product clusters to be analyzed increases. When the string similarity coefficient increases, the similarity feature index also increases, and the similarity between the names of the products to be analyzed and between the product clusters to be analyzed also increases; a more accurate assessment of the similarity between the names of the products to be analyzed and between the product clusters to be analyzed is achieved, further improving the accuracy of the name classification analysis of manufacturing products.

[0052] Furthermore, a specific method for obtaining cluster analysis results is as follows: defining the product name feature vector as an initial cluster, merging the initial clusters whose similarity feature indexes reach a preset clustering point into a first cluster; updating the similarity matrix based on the first cluster, the similarity matrix is ​​used to reflect the degree of similarity between the first cluster and the remaining initial clusters; repeating the above steps until all initial clusters are merged into a hierarchical clustering tree, and the hierarchical clustering tree represents the cluster analysis results.

[0053] In this embodiment, the preset clustering point is the minimum similarity index set by the preset personnel, and the minimum similarity index is as low as 0.9; the accurate acquisition of clustering analysis results is achieved, and the accuracy of name classification analysis of manufacturing products is further improved.

[0054] Furthermore, the method for obtaining the quality assessment index is as follows: obtain the cluster quality score according to the cluster analysis results, the cluster quality score includes the intra-cluster similarity and the inter-cluster offset; obtain the central cluster micro-number according to the obtained cluster quality score, the central cluster micro-number includes the central similarity micro-number and the central offset micro-number; obtain the quality assessment index of the cluster analysis according to the obtained cluster quality score and the central cluster micro-number.

[0055] In this embodiment, the quality evaluation index is calculated using the following formula:

[0056]

[0057] Where t represents the number of cluster analysis, t = 1, 2, ..., T, T is the total number of cluster analysis, r represents the number of clusters in the cluster analysis results, r = 1, 2, ..., R, R is the total number of clusters in the cluster analysis results, ZLt represents the quality evaluation index of the t-th cluster analysis, represents the intra-cluster similarity of clusters in the r-th cluster analysis result of the t-th cluster analysis, It represents the center similarity of the clusters in the rth cluster analysis result of the tth cluster analysis. Indicates the inter-cluster offset of clusters in the r-th cluster analysis result of the t-th cluster analysis, The micronumber representing the center offset of the cluster in the r-th cluster analysis result of the t-th cluster analysis.

[0058] Specifically, the intra-cluster similarity is the similarity between the names of the products to be analyzed within the cluster, the inter-cluster offset is the difference between different clusters, the similarity between the names of the products to be analyzed within the cluster is measured by the average intra-cluster distance and the average distance of the point pairs within the cluster, the center similarity micronumber represents the similarity with the center cluster, and the center offset micronumber represents the offset with the center cluster; the difference between clusters is measured by the distance between the center points of different clusters; the quality of the cluster analysis of the product name feature vector is quantified, which further improves the accuracy of the name classification analysis of manufacturing products.

[0059] Furthermore, after obtaining the first cluster analysis result, the first clustering result is also visualized, and the visualization content is as follows: a cluster diagram is drawn after the first cluster analysis result is reduced in dimension, and the cluster diagram includes a cluster scatter diagram, a cluster heat map and a cluster parallel coordinate diagram; cluster annotations are made in the cluster diagram, and according to the visualization content, the relationship between different clusters and the distribution of the clusters are monitored in real time.

[0060] In this embodiment, the cluster scatter plot is used to display the name numerical vector to be analyzed in the first cluster analysis result in two-dimensional and three-dimensional space, and to mark clusters at different levels; the cluster heat map is used to display the tree structure and cluster merging in the hierarchical clustering tree, and use color depth to represent the similarity and distance between the name numerical vector to be analyzed and between clusters; the cluster parallel coordinate diagram plots the product name feature vectors of the name numerical vector to be analyzed on vertical parallel axes, and different clusters are represented by different colors or lines to show the distribution and clustering of data in high-dimensional space; the cluster annotation includes cluster center, cluster boundary and data point label; the visualization of cluster analysis of products to be analyzed is realized, and the accuracy of name classification analysis of manufacturing products is further improved.

[0061] To summarize, the embodiment of the present application obtains the naming pattern of the product name to be analyzed by constructing a naming pattern recognition model, and converts the product name to be analyzed obtained in real time into a numerical vector of the name to be analyzed, and then obtains the similarity feature index based on the extracted product name feature vector, and then obtains the cluster analysis result based on the product name feature vector, and finally obtains the quality assessment index for the cluster analysis result, and obtains the first cluster analysis result based on the quality assessment index, thereby realizing the cluster analysis of the names of manufacturing products, and further improving the accuracy of the classification analysis of the names of manufacturing products, and effectively solving the problem of low accuracy of the classification analysis of the names of manufacturing products in the prior art.

[0062] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0063] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0064] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0066] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0067] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A product clustering analysis method based on a deep learning model, characterized in that: The following steps are involved: S1, collecting product names of manufacturing products from a product name database and constructing a naming pattern recognition model, wherein the naming pattern recognition model is used to recognize the naming pattern of the product names to be analyzed obtained in real time; S2, combining the naming pattern recognition model to obtain the naming pattern of the product name to be analyzed, and converting the real-time obtained product name to be analyzed into a numerical vector of the name to be analyzed; S3, combining the result of semantic analysis of the name numerical vector to be analyzed with the product name feature vector of the name numerical vector to be analyzed extracted by the clustering target to obtain a similarity feature index, and obtaining a cluster analysis result according to the product name feature vector, wherein the similarity feature index is used to describe the similarity between the names of the products to be analyzed and between the clusters of the products to be analyzed; S4, performing a quality assessment on the clustering analysis result to obtain a quality assessment index, optimizing the clustering algorithm according to the quality assessment index, and then repeatedly performing S1 until the quality assessment index meets a corresponding threshold and obtains a first clustering analysis result, wherein the quality assessment index is used to describe the quality of the clustering analysis of the product name feature vector; The method for obtaining the similarity feature index is as follows: Obtain similarity feature scores between the product names to be analyzed based on the extracted product name feature vectors, wherein the similarity feature scores include semantic similarity and string similarity; Obtaining a similar feature score vector according to the obtained similar feature scores, and simultaneously obtaining a corresponding reference vector, wherein the similar feature score vector includes a semantic similarity vector and a string similarity vector; The similarity feature index of the product to be analyzed is obtained based on the obtained similarity feature score vector and its corresponding reference vector. The similarity feature index is calculated using the following formula: In the formula, m represents the number of the product name to be analyzed, m = 1, 2, ..., M, M is the total number of product names to be analyzed, XS m represents the similarity feature index of the mth product name to be analyzed, represents the semantic similarity vector of the mth product name to be analyzed, represents the reference semantic similarity vector of the mth product name to be analyzed, represents the string similarity vector of the mth product name to be analyzed, The reference string similarity vector representing the name of the mth product to be analyzed; The specific method for obtaining cluster analysis results is as follows: The product name feature vector is defined as the initial cluster, and the initial clusters whose similar feature indexes reach the preset aggregation point are merged into the first cluster; updating a similarity matrix according to the first cluster, wherein the similarity matrix is ​​used to reflect the similarity between the first cluster and the remaining initial clusters; The above steps are repeated until all initial clusters are merged into a hierarchical clustering tree, which represents the clustering analysis result.

2. The product clustering analysis method based on the deep learning model according to claim 1, characterized in that: The steps for constructing the naming pattern recognition model are: S11, dividing the product names after the named entity annotation into an entity training set, an entity verification set, and an entity test set according to a preset ratio; S12, according to the selected hyperparameters and training parameters of the recurrent neural network model, and using the entity training set to train the recurrent neural network model until the model converges to obtain a training model; S13, use the entity verification set to perform performance verification on the training model to obtain performance verification results, and determine whether the performance verification results meet the corresponding threshold conditions. If so, generate a naming pattern recognition model and input the named entity to recognize the naming pattern. Otherwise, re-train the model until the performance verification results meet the corresponding threshold conditions.

3. The product clustering analysis method based on the deep learning model as claimed in claim 2, characterized in that: The specific steps of identifying the naming pattern are: Performing model testing on the training model using the entity test set to obtain an entity test index, where the entity test index is used to test the performance of the training model; Determine whether the obtained entity test index meets the test threshold. If so, identify the named entity through the training model. If not, optimize the parameters of the training model. The identified named entities and the corresponding named entity types are output to obtain the identified naming patterns.

4. The product clustering analysis method based on the deep learning model as claimed in claim 3, characterized in that: The method for obtaining the entity test index is as follows: Extracting an entity test score from the model testing process, the entity test score includes a learning rate, an accuracy rate, and a recall rate, and obtaining a corresponding preset standard learning rate according to the learning rate; Obtaining an entity test index of the training model according to the obtained entity test score and a preset standard learning rate; The entity test index is calculated using the following formula: Where e is a natural constant, g is the number of the training round, g = 1, 2, ..., G, G is the total number of training rounds, ST represents the entity test index of the training model, X g represents the learning rate of the training model in the gth training round, α g represents the accuracy of the training model in the gth training round, β g represents the recall rate of the training model in the gth training round, and X0 represents the preset standard learning rate of the training model.

5. The product clustering analysis method based on the deep learning model according to claim 1, characterized in that: The process of converting the real-time acquired product names to be analyzed into the name numerical vector to be analyzed is as follows: Delete redundant symbols in the names of products to be analyzed obtained in real time, and convert all names of products to be analyzed into lowercase; Use the word segmentation tool to split the product name to be analyzed into separate words, and use the stem extraction method to restore the separate words until they become the root form of the word; Get the vocabulary prototype set and convert it into a bag of words; The bag-of-words model is combined to map the product name to be analyzed with the numerical vector and the generated numerical vector is standardized to obtain the numerical vector of the name to be analyzed.

6. The product clustering analysis method based on the deep learning model according to claim 1, characterized in that: The method for obtaining the quality assessment index is as follows: Obtaining a cluster quality score according to the cluster analysis result, wherein the cluster quality score includes intra-cluster similarity and inter-cluster offset; Obtaining a central cluster micro-number according to the obtained cluster quality score, wherein the central cluster micro-number includes a central similarity micro-number and a central offset micro-number; The quality assessment index of cluster analysis is obtained based on the obtained cluster quality score and central cluster number.

7. The product clustering analysis method based on the deep learning model according to claim 1, characterized in that: The first cluster analysis result is obtained, and then the first cluster result is visualized, and the content of the visualization is as follows: Draw a cluster diagram after performing dimensionality reduction on the first cluster analysis result, wherein the cluster diagram includes a cluster scatter diagram, a cluster heat map and a cluster parallel coordinate diagram; Annotate clusters in the cluster diagram and monitor the relationships between different clusters and their distribution in real time based on the visualization content.

Citation Information

Patent Citations

  • Parallelization of massive data clustering analysis

    CN102855259B

  • A Fast Association Method of Objects Based on Cluster Analysis

    CN103617163B

  • Social media short text online clustering method based on entity constraint

    CN110442726A

  • Incremental clustering algorithm based on community detection

    CN110990566A