Model determination method and device, equipment and computer storage medium

By selecting different clustering models to form the target clustering model and using the objective function value for filtering, the problems of low clustering efficiency and insufficient accuracy in existing technologies are solved, achieving efficient and accurate clustering of user experience feedback information, which is applicable to a variety of business scenarios.

CN117171347BActive Publication Date: 2026-03-24CCB FINTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-19
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In existing technologies, the clustering methods for user experience feedback information are singular, resulting in low clustering efficiency and inaccurate results, and strong limitations in applicable business scenarios.

Method used

By selecting different clustering models to form the final target clustering model, and using the objective function value of the clustering model, the smallest clustering model is selected and added to the set until the preset number is reached, forming a multi-model combination for cluster analysis.

Benefits of technology

It improves clustering efficiency and accuracy, is applicable to most business scenarios, and has universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117171347B_ABST
    Figure CN117171347B_ABST
Patent Text Reader

Abstract

The application discloses a model determination method and device, equipment and a computer storage medium. The method comprises the following steps: obtaining a first sample set and a preset clustering model, the first sample set comprising second preset quantity and preset dimension of vectorized data; clustering the first sample set by using the preset clustering model to obtain a clustering result, and calculating a first target function value of the preset clustering model; adding a first clustering model with the minimum first target function value to a clustering model set, selecting a second clustering model from the preset clustering model, and adding the second clustering model to the clustering model set; selecting a third clustering model from the preset clustering model, adding the third clustering model to the clustering model set, until the number of clustering models in the clustering model set is not less than a third preset number, and determining all the clustering models in the clustering model set as target clustering models. The efficiency of clustering and the accuracy of the clustering result are improved, the method is suitable for most business scenarios, and is universal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular to a method, apparatus, device, and computer storage medium for determining a model. Background Technology

[0002] When optimizing business products, it is essential to perform cluster analysis on user experience feedback information.

[0003] In existing technologies, after collecting user experience feedback, the feedback is typically clustered through manual screening, or designers set fixed predictions based on past experience and then cluster the feedback based on those predictions. However, clustering methods or models based on manual screening or fixed predictions are relatively simplistic, leading to low efficiency, inaccurate results, and limitations in applicable business scenarios. Summary of the Invention

[0004] This application provides a method, apparatus, device, and computer storage medium for determining a model. By selecting different clustering models to form the final target clustering model through the objective function value of the clustering model, the method avoids the use of a single clustering model, which not only improves the efficiency and accuracy of clustering results, but also eliminates the limitations of business scenarios and is applicable to most business scenarios, thus possessing versatility.

[0005] In a first aspect, embodiments of this application provide a method for determining a model, including:

[0006] Obtain a first sample set of a first preset number and a preset clustering model. The first sample set includes vectorized data of a second preset number and preset dimensions.

[0007] The first sample set is clustered using a pre-defined clustering model to obtain clustering results. The first objective function value of the pre-defined clustering model is calculated, and the first objective function value characterizes the degree of clustering of the clustering results.

[0008] The first clustering model with the smallest first objective function value is added to the clustering model set. The second clustering model is selected from the preset clustering models and added to the clustering model set.

[0009] If the number of cluster models in the cluster model set is less than the third preset number, select the third cluster model from the preset cluster models and add the third cluster model to the cluster model set until the number of cluster models in the cluster model set is not less than the third preset number. Then, determine all cluster models in the cluster model set as the target cluster model.

[0010] Among them, the second objective function values ​​of the first clustering model and the second clustering model are less than the first objective function value of the first clustering model, and the third objective function values ​​of the first clustering model, the second clustering model and the third clustering model are less than the second objective function value.

[0011] In one possible implementation embodiment, it further includes:

[0012] Based on the clustering results of each clustering model in the target clustering model, a third preset number of adjacency matrices are generated;

[0013] Calculate the weighted average of the adjacency matrices of the third preset number;

[0014] The category of each sample in the first sample set of the first preset number is determined based on the weighted average.

[0015] In one possible implementation, before obtaining a first preset number of first sample sets and a preset clustering model, the method further includes:

[0016] Obtain text data from multiple dimensions;

[0017] The text data is converted into vector data to obtain the second sample set;

[0018] Randomly select a second preset number and preset dimension of vectorized data from the second sample set to obtain a first sample set of a first preset number.

[0019] In one possible implementation, the preset clustering model includes a first sub-clustering model, and the method further includes:

[0020] In the absence of label information in the first sample set, the first sample set is clustered using the first sub-clustering model to obtain the first clustering result, and the fourth objective function value of the first preset clustering model is calculated using the first objective function; wherein, the label information includes information that the first sample and the second sample in the first sample set belong to the same category;

[0021] The fourth clustering model with the smallest fourth objective function value is added to the first clustering model set. The fifth clustering model is selected from the first sub-clustering models and added to the first clustering model set.

[0022] If the number of cluster models in the first cluster model set is less than the third preset number, select the sixth cluster model from the first sub-cluster models and add the sixth cluster model to the first cluster model aggregation until the number of cluster models in the first cluster model set is not less than the third preset number, and determine all cluster models in the first cluster model set as the first target cluster model;

[0023] Among them, the fifth objective function values ​​of the fourth and fifth clustering models are less than the fourth objective function value of the fourth clustering model, and the sixth objective function values ​​of the fourth, fifth, and sixth clustering models are less than the fifth objective function value.

[0024] In one possible implementation, the preset clustering model includes a second sub-clustering model, and the method further includes:

[0025] When the first sample set includes labeling information, the first sample set is clustered using a two-sub-clustering model to obtain a second clustering result. Then, the seventh objective function value of the second preset clustering model is calculated using a second objective function. The labeling information includes information that the first sample and the second sample in the first sample set belong to the same category.

[0026] The seventh clustering model with the smallest objective function value is added to the second clustering model set. The eighth clustering model is selected from the second sub-clustering models and added to the second clustering model set.

[0027] If the number of cluster models in the second cluster model set is less than the third preset number, select the ninth cluster model from the second sub-cluster models and add the ninth cluster model to the second cluster model set until the number of cluster models in the second cluster model set is not less than the third preset number. Then, determine all cluster models in the second cluster model set as the second target cluster model.

[0028] Among them, the eighth objective function value of the seventh clustering model and the eighth clustering model is less than the seventh objective function value of the seventh clustering model, and the ninth objective function value of the seventh clustering model, the eighth clustering model and the ninth clustering model is less than the eighth objective function value.

[0029] In one possible implementation, the first objective function satisfies the following condition:

[0030]

[0031] in, d(p) represents the cluster center of category h. i ,μ h ) represents sample p i The Euclidean distance between the cluster centers and h represents the category obtained through each preset clustering model, θ is the indicator function, and θ(y) is the distance between the cluster centers and h. i =h) is used to determine sample p i Category y i Is it h? If not, it is 0. P is the first sample set, and k represents the number of sample categories.

[0032] In one possible implementation, the second objective function satisfies the following condition:

[0033]

[0034] in, d(p) represents the cluster center of category h. i μ h ) represents sample p i The Euclidean distance between the cluster centers and h represents the category obtained through each preset clustering model, θ is the indicator function, and θ(y) is the distance between the cluster centers and h. i =h) is used to determine sample p i Category y i Is it h? If not, then it is 0. P is the first sample set. Characterizing the actual sample p i Category y i and sample p j Category y j The penalty parameter θ(y) for classifying the same items as different categories. i ≠y j Used to determine sample p i Category y i and sample p j Category y j If they are the same, the value is 0. Characterizing the actual sample p i Category y i and sample p j Category y j The penalty parameter θ(y) for classifying different items as belonging to the same category. i =y j Used to determine sample p i Category y i and sample p j Category y j Whether they are the same or different, the value is 0 if they are different. M and N identify the first sample set, and k represents the number of sample categories.

[0035] Secondly, embodiments of this application provide a model determining apparatus, comprising:

[0036] The acquisition module is used to acquire a first sample set of a first preset number and a preset clustering model. The first sample set includes vectorized data of a second preset number and preset dimensions.

[0037] The determination module is used to cluster the first sample set using a preset clustering model, obtain the clustering results, and calculate the first objective function value of the preset clustering model. The first objective function value characterizes the degree of clustering of the clustering results.

[0038] An addition module is used to add the first clustering model with the smallest first objective function value to the clustering model set, select a second clustering model from the preset clustering models, and add the second clustering model to the clustering model set;

[0039] The addition module is also used to select a third clustering model from the preset clustering models when the number of clustering models in the clustering model set is less than the third preset number, and add the third clustering model to the clustering model set until the number of clustering models in the clustering model set is not less than the third preset number, and determine all clustering models in the clustering model set as the target clustering model;

[0040] Among them, the second objective function values ​​of the first clustering model and the second clustering model are less than the first objective function value of the first clustering model, and the third objective function values ​​of the first clustering model, the second clustering model and the third clustering model are less than the second objective function value.

[0041] Thirdly, embodiments of this application provide an electronic device, the device comprising:

[0042] Processor and memory storing computer program instructions;

[0043] The method for determining the model in which the processor implements any of the above when executing computer program instructions.

[0044] Fourthly, embodiments of this application provide a computer storage medium on which computer program instructions are stored, and a method for determining the model that implements any of the above-mentioned items when the computer program instructions are executed by a processor.

[0045] Fifthly, embodiments of this application provide a computer program product, characterized in that, when the instructions in the computer program product are executed by the processor of an electronic device, a method for determining the model that enables the electronic device to execute any of the above-mentioned items is provided.

[0046] This application discloses a method, apparatus, device, and computer storage medium for determining a model. The method includes: acquiring a first sample set of a first preset number and a preset clustering model, wherein the first sample set includes vectorized data of a second preset number and a preset dimension; clustering the first sample set using the preset clustering model to obtain clustering results, and calculating a first objective function value of the preset clustering model, wherein the first objective function value characterizes the degree of clustering of the clustering results; adding the first clustering model with the smallest first objective function value to the clustering model set; selecting a second clustering model from the preset clustering models; and adding the second clustering model to the set. The process involves selecting a set of clustering models. If the number of clustering models in the set is less than a third preset number, a third clustering model is selected from the preset models and added to the set. This process continues until the number of clustering models in the set is not less than the third preset number. All clustering models in the set are then considered the target clustering model. Specifically, the second objective function values ​​of the first and second clustering models are less than the first objective function value of the first clustering model, and the third objective function values ​​of the first, second, and third clustering models are less than the second objective function value. This method selects different clustering models based on their objective function values ​​to form the final target clustering model, avoiding the use of a single clustering model. This not only improves the efficiency and accuracy of clustering results but also overcomes the limitations of specific business scenarios, making it applicable to most business scenarios and possessing versatility. Attached Figure Description

[0047] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart illustrating a method for determining a model according to an embodiment of this application;

[0049] Figure 2 This is a flowchart illustrating a method for determining a model provided in another embodiment of this application;

[0050] Figure 3 This is a schematic diagram of the structure of a model determining device provided in another embodiment of this application;

[0051] Figure 4 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0052] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0053] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0054] It should be noted that the acquisition, storage, use, and processing of data in this application embodiment all comply with the relevant provisions of national laws and regulations.

[0055] When optimizing business products, it is essential to perform cluster analysis on user experience feedback information.

[0056] In existing technologies, after collecting user experience feedback, the feedback is typically clustered through manual screening, or designers set fixed predictions based on past experience and then cluster the feedback based on those predictions. However, clustering methods or models based on manual screening or fixed predictions are relatively simplistic, leading to low efficiency, inaccurate results, and limitations in applicable business scenarios.

[0057] To address the problems of the prior art, embodiments of this application provide a method, apparatus, device, and computer storage medium for determining a model. The method for determining a model provided in this application will be described first.

[0058] Figure 1 A flowchart illustrating a method for determining a model according to an embodiment of this application is shown.

[0059] like Figure 1As shown, the model determination method provided in this application embodiment includes the following S110 to S140.

[0060] S110. Obtain a first sample set of a first preset quantity and a preset clustering model. The first sample set includes vectorized data of a second preset quantity and preset dimensions.

[0061] Here, the first preset quantity is set in advance, the second preset quantity is set in advance, and the preset dimension is set in advance.

[0062] In some embodiments, different first sample sets may include the same samples.

[0063] In some embodiments, vectorized data can be recognized by a clustering model. The preset clustering model includes at least one, which can be an unsupervised clustering model or a semi-supervised clustering model. For example, unsupervised clustering models may include, but are not limited to, K-means clustering and DBSCAN clustering models, and semi-supervised clustering models may include, but are not limited to, PCKmeans clustering and E2CP clustering models.

[0064] As an example, each first sample set consists of 100 samples across 10 dimensions.

[0065] S120. Cluster the first sample set using a preset clustering model to obtain the clustering results, and calculate the first objective function value of the preset clustering model. The first objective function value represents the degree of clustering of the clustering results.

[0066] Here, the value of the first objective function characterizes the degree of clustering of the clustering results. The smaller the value of the first objective function, the more compact the clustering results.

[0067] In some embodiments, the clustering result is the clustering result of samples in the first sample set. For different first sample sets, it may be the clustering result of the same sample.

[0068] In some embodiments, for each first sample set, a preset clustering model is used to cluster each first sample set, the first objective function value of each preset clustering model is calculated, and a preset clustering model is selected based on different first objective function values.

[0069] In some embodiments, the values ​​of the first objective function can be sorted from smallest to largest to obtain a sorting result. From the sorting result, the preset clustering model corresponding to the first first objective function value is selected as the first clustering model.

[0070] In some embodiments, two first objective function values ​​can be arbitrarily selected, their magnitudes compared, and the smaller value retained. Then, a first objective function value other than the two previously selected values ​​is chosen and compared with the retained value; the smaller value is retained. This selection and comparison process is repeated until the smallest first objective function value is selected, and the preset clustering model with the smallest first objective function value is adopted as the first clustering model.

[0071] S130. Add the first clustering model with the smallest first objective function value to the clustering model set. Select the second clustering model from the preset clustering models and add the second clustering model to the clustering model set.

[0072] Here, the first clustering model is the one with the smallest first objective function value among the preset clustering models. The second clustering model is selected such that the second objective function values ​​of the first and second clustering models are less than the first objective function value of the first clustering model.

[0073] In some embodiments, after adding the first clustering model to the clustering model set, a second clustering model is selected from the preset clustering models. The second objective function values ​​of the first and second clustering models are calculated, and the magnitudes of the first and second objective function values ​​are compared. If the second objective function value is less than the first objective function value, the selected second clustering model is added to the clustering model set; if the second objective function value is not less than the first objective function value, a new second clustering model is selected. It is understood that the second clustering model is different from the first clustering model.

[0074] S140. If the number of cluster models in the cluster model set is less than the third preset number, select the third cluster model from the preset cluster models and add the third cluster model to the cluster model set until the number of cluster models in the cluster model set is not less than the third preset number. Then, determine all cluster models in the cluster model set as target cluster models. Among them, the second objective function value of the first cluster model and the second cluster model is less than the first objective function value of the first cluster model, and the third objective function value of the first cluster model, the second cluster model and the third cluster model is less than the second objective function value.

[0075] Here, the third preset quantity is predetermined. The selection of the third clustering model satisfies the condition that the third objective function value of the first, second, and third clustering models is less than the second objective function value.

[0076] In some embodiments, after adding the second clustering model to the clustering model set, it is determined whether the number of clustering models included in the clustering model set is less than a third preset number. If it is less, a third clustering model is selected from the preset clustering models, and the third objective function values ​​of the first, second, and third clustering models are calculated. The second and third objective function values ​​are then compared. If the third objective function value is less than the second objective function value, the selected third clustering model is added to the clustering model set; if the third objective function value is not less than the second objective function value, a new third clustering model is selected. It is understood that the third clustering model is different from the first and second clustering models. If the number of clustering models included in the clustering model set is not less than the third preset number, the first and second clustering models are determined as target clustering models. It is understood that when the first and second clustering models are determined as target clustering models, the third preset number is 2.

[0077] In some embodiments, after adding the third clustering model to the clustering model set, it is determined whether the number of clustering models in the clustering model set is less than a third preset number. If it is less, clustering models are selected from the preset clustering models and added to the clustering model set. After adding more clustering models, the objective function values ​​of all clustering models in the clustering model set are always less than (or not greater than) the objective function values ​​of all clustering models in the clustering model set before adding the new models. This process continues until the number of clustering models in the clustering model set is not less than the third preset number, at which point all clustering models in the clustering model set are determined as target clustering models. If the number of clustering models in the clustering model set is not less than the third preset number, then the first, second, and third clustering models are determined as target clustering models. It can be understood that when the first, second, and third clustering models are determined as target clustering models, the third preset number is 3.

[0078] In this way, by selecting different clustering models to form the final target clustering model through the objective function value of the clustering model, the single clustering model is avoided. This not only improves the efficiency and accuracy of clustering results, but also breaks free from the limitations of business scenarios, making it applicable to most business scenarios and possessing universality.

[0079] Based on this, in some embodiments, the method may further include:

[0080] Based on the clustering results of each clustering model in the target clustering model, a third preset number of adjacency matrices are generated;

[0081] Calculate the weighted average of the adjacency matrices of the third preset number;

[0082] The category of each sample in the first sample set of the first preset number is determined based on the weighted average.

[0083] In some embodiments, the weighted average of the adjacency matrix is ​​calculated using formula (1), which is as follows:

[0084]

[0085] Where O represents the weighted average of the adjacency matrices, and B represents the number of adjacency matrices. b This represents the product of the b-th adjacency matrix and its corresponding weight.

[0086] In some embodiments, the clustering consensus function of the Normalized Cuts algorithm is used to determine the category of each sample in a first sample set of a first preset number of samples based on a weighted average, and the clustering result of each sample is output. It should be noted that the consensus function combines multiple clustering results in a cluster set to generate a unified clustering result. The clustering result includes the category of the samples, and the adjacency matrix is ​​constructed as follows: if the clustering results of two samples are in the same category, they are considered adjacent.

[0087] As an example, categories may include, but are not limited to, UI, information, interaction, architecture, process, functionality, and performance.

[0088] In this way, clustering the first sample set using multiple clustering models improves clustering performance and yields more accurate clustering results.

[0089] Based on this, in some embodiments, the method may further include S101 to S103 before S110 described above.

[0090] S101. Obtain text data from multiple dimensions.

[0091] In some embodiments, the text data can be user experience feedback information. Here, the text data can be multi-dimensional.

[0092] S102. Convert the text data into vectorized data to obtain the second sample set.

[0093] In some embodiments, text data can be transformed into vectorized data through statistical methods or neural networks, without specific limitations. For example, statistical methods include the bag-of-words model and the term frequency–inverse document frequency (TF-IDF) model; neural network-based methods include the word2vec model (a related model used to generate word vectors), the ELMo model, and the BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language representation model) model.

[0094] S103. Randomly select vectorized data of a second preset quantity and preset dimension from the second sample set to obtain a first sample set of a first preset quantity.

[0095] In some embodiments, vectorized data are randomly selected from the second sample set according to a second preset quantity and a preset dimension to obtain a first sample set of a first preset quantity. Different first sample sets may include the same samples.

[0096] In this way, text data is transformed into vectorized data that the clustering model can recognize, so that the clustering model can cluster the samples and obtain the clustering results.

[0097] Based on this, in some embodiments, the preset clustering model includes a first sub-clustering model, and the method may further include:

[0098] In the absence of label information in the first sample set, the first sample set is clustered using the first sub-clustering model to obtain the first clustering result, and the fourth objective function value of the first preset clustering model is calculated using the first objective function; wherein, the label information includes information that the first sample and the second sample in the first sample set belong to the same category;

[0099] The fourth clustering model with the smallest fourth objective function value is added to the first clustering model set. The fifth clustering model is selected from the first sub-clustering models and added to the first clustering model set.

[0100] If the number of cluster models in the first cluster model set is less than the third preset number, select the sixth cluster model from the first sub-cluster models and add the sixth cluster model to the first cluster model aggregation until the number of cluster models in the first cluster model set is not less than the third preset number, and determine all cluster models in the first cluster model set as the first target cluster model;

[0101] Among them, the fifth objective function values ​​of the fourth and fifth clustering models are less than the fourth objective function value of the fourth clustering model, and the sixth objective function values ​​of the fourth, fifth, and sixth clustering models are less than the fifth objective function value.

[0102] Here, the first sub-clustering model includes the unsupervised clustering model. The first objective function only needs to reflect the clustering degree of the first sub-clustering model, and no specific limitation is made here. It can be understood that, for the first sample set, if the information that the first and second samples in the first sample set belong to the same category is not included, then the value of the fourth objective function is only related to the distance between the two samples. The smaller the distance, the smaller the value of the fourth objective function, and the more compact the clustering result.

[0103] In this way, the target clustering model can be determined when the first sample set does not include label information.

[0104] Based on this, in some embodiments, the preset clustering model includes a second sub-clustering model, and the method may further include:

[0105] When the first sample set includes labeling information, the first sample set is clustered using a two-sub-clustering model to obtain a second clustering result. Then, the seventh objective function value of the second preset clustering model is calculated using a second objective function. The labeling information includes information that the first sample and the second sample in the first sample set belong to the same category.

[0106] The seventh clustering model with the smallest objective function value is added to the second clustering model set. The eighth clustering model is selected from the second sub-clustering models and added to the second clustering model set.

[0107] If the number of cluster models in the second cluster model set is less than the third preset number, select the ninth cluster model from the second sub-cluster models and add the ninth cluster model to the second cluster model set until the number of cluster models in the second cluster model set is not less than the third preset number. Then, determine all cluster models in the second cluster model set as the second target cluster model.

[0108] Among them, the eighth objective function value of the seventh clustering model and the eighth clustering model is less than the seventh objective function value of the seventh clustering model, and the ninth objective function value of the seventh clustering model, the eighth clustering model and the ninth clustering model is less than the eighth objective function value.

[0109] Here, the second sub-clustering model includes a semi-supervised clustering model. The second objective function only needs to reflect the clustering degree of the second sub-clustering model, and is not specifically limited here. It can be understood that, for the first sample set, if it includes information that the first and second samples in the first sample set belong to the same category, then the value of the seventh objective function is related not only to the distance between the two samples, but also to the accuracy of the clustering result. The smaller the distance, the more accurate the clustering result; the smaller the value of the seventh objective function, the more compact the clustering result.

[0110] In this way, the target clustering model can be determined when the first sample set includes label information.

[0111] Based on this, in some embodiments, the first objective function can satisfy formula (2), which is as follows:

[0112]

[0113] in, d(p) represents the cluster center of category h. i ,μ h ) represents sample p i The Euclidean distance between the cluster centers and h represents the category obtained through each preset clustering model, θ is the indicator function, and θ(y) is the distance between the cluster centers and h. i =h) is used to determine sample p i Category y i Is it h? If not, then it is 0. P is the first sample set, k represents the number of sample categories, and Δ(I) represents the value of the fourth objective function.

[0114] In this way, the distance between samples can be calculated using the first objective function, and the smaller the value of the fourth objective function, the more compact the clustering result.

[0115] Based on this, in some embodiments, the second objective function can satisfy formula (3), which is as follows:

[0116]

[0117] in, d(p) represents the cluster center of category h. i ,μ h ) represents sample p i The Euclidean distance between the cluster centers and h represents the category obtained through each preset clustering model, θ is the indicator function, and θ(y) is the distance between the cluster centers and h. i =h) is used to determine sample p i Category y i Is it h? If not, then it is 0. P is the first sample set. Characterizing the actual sample p i Category yi and sample p j Category y j The penalty parameter θ(y) for classifying the same items as different categories. i ≠y j Used to determine sample p i Category y i and sample p j Category y j If they are the same, the value is 0. Characterizing the actual sample p i Category y i and sample p j Category y j The penalty parameter θ(y) for classifying different items as belonging to the same category. i =y j Used to determine sample p i Category y i and sample p j Category y j Whether they are the same or different, if they are determined to be different, then the value is 0. M and N identify the first sample set, k represents the number of sample categories, and Δ(I) represents the value of the seventh objective function.

[0118] In this way, the distance between samples can be calculated using the second objective function, and the smaller the value of the seventh objective function, the more compact the clustering result.

[0119] Based on the model determination method provided in the above embodiments, this application also provides specific implementations of the model determination apparatus. Please refer to the following embodiments.

[0120] See Figure 3 The model determining device 300 provided in this application embodiment includes:

[0121] The acquisition module 310 is used to acquire a first sample set of a first preset number and a preset clustering model. The first sample set includes vectorized data of a second preset number and preset dimensions.

[0122] The determination module 320 is used to cluster the first sample set using a preset clustering model, obtain the clustering results, and calculate the first objective function value of the preset clustering model. The first objective function value represents the degree of clustering of the clustering results.

[0123] Add module 330 to add the first clustering model with the smallest first objective function value to the clustering model set, select a second clustering model from the preset clustering models, and add the second clustering model to the clustering model set;

[0124] The added module 330 is also used to select a third clustering model from the preset clustering models when the number of clustering models in the clustering model set is less than the third preset number, and add the third clustering model to the clustering model set until the number of clustering models in the clustering model set is not less than the third preset number, and determine all clustering models in the clustering model set as the target clustering model;

[0125] Among them, the second objective function values ​​of the first clustering model and the second clustering model are less than the first objective function value of the first clustering model, and the third objective function values ​​of the first clustering model, the second clustering model and the third clustering model are less than the second objective function value.

[0126] Based on this, in some embodiments, the device 300 may further include:

[0127] Based on the clustering results of each clustering model in the target clustering model, a third preset number of adjacency matrices are generated;

[0128] Calculate the weighted average of the adjacency matrices of the third preset number;

[0129] The category of each sample in the first sample set of the first preset number is determined based on the weighted average.

[0130] Based on this, in some embodiments, the device 300 may further include:

[0131] The acquisition module 310 is also used to acquire text data of multiple dimensions before acquiring the first sample set of the first preset number and the preset clustering model;

[0132] The transformation module is used to convert text data into vectorized data to obtain the second sample set.

[0133] The selection module is used to randomly select a second preset number and preset dimension of vectorized data from the second sample set to obtain a first sample set of a first preset number.

[0134] In one possible implementation, the preset clustering model includes a first sub-clustering model, and the device 300 may further include:

[0135] The determining module 320 is further configured to, when the first sample set does not include labeling information, cluster the first sample set using a first sub-clustering model to obtain a first clustering result, and use a first objective function to calculate the fourth objective function value of the first preset clustering model; wherein, the labeling information includes information that the first sample and the second sample in the first sample set belong to the same category;

[0136] The added module 330 is also used to add the fourth clustering model with the smallest fourth objective function value to the first clustering model set, select the fifth clustering model from the first sub-clustering model, and add the fifth clustering model to the first clustering model set;

[0137] The added module 330 is also used to select a sixth clustering model from the first sub-clustering model when the number of clustering models in the first clustering model set is less than the third preset number, and add the sixth clustering model to the first clustering model aggregation until the number of clustering models in the first clustering model set is not less than the third preset number, and determine all clustering models in the first clustering model set as the first target clustering model;

[0138] Among them, the fifth objective function values ​​of the fourth and fifth clustering models are less than the fourth objective function value of the fourth clustering model, and the sixth objective function values ​​of the fourth, fifth, and sixth clustering models are less than the fifth objective function value.

[0139] Based on this, in some embodiments, the preset clustering model includes a second sub-clustering model, and the device 300 may further include:

[0140] The determining module 320 is further configured to, when the first sample set includes labeling information, use a two-sub-clustering model to cluster the first sample set to obtain a second clustering result, and use a second objective function to calculate the seventh objective function value of the second preset clustering model; wherein, the labeling information includes information that the first sample and the second sample in the first sample set belong to the same category;

[0141] The added module 330 is also used to add the seventh clustering model with the smallest seventh objective function value to the second clustering model set, select the eighth clustering model from the second sub-clustering models, and add the eighth clustering model to the second clustering model set;

[0142] The added module 330 is also used to select a ninth clustering model from the second sub-clustering models when the number of clustering models in the second clustering model set is less than the third preset number, and add the ninth clustering model to the second clustering model set until the number of clustering models in the second clustering model set is not less than the third preset number, and determine all clustering models in the second clustering model set as the second target clustering model.

[0143] Among them, the eighth objective function value of the seventh clustering model and the eighth clustering model is less than the seventh objective function value of the seventh clustering model, and the ninth objective function value of the seventh clustering model, the eighth clustering model and the ninth clustering model is less than the eighth objective function value.

[0144] Based on this, in some embodiments, the first objective function satisfies the following condition:

[0145]

[0146] in, d(p) represents the cluster center of category h. i ,μ h ) represents sample p i The Euclidean distance between the cluster centers and h represents the category obtained through each preset clustering model, θ is the indicator function, and θ(y) is the distance between the cluster centers and h. i =h) is used to determine sample p i Category y i Is it h? If not, it is 0. P is the first sample set, and k represents the number of sample categories.

[0147] Based on this, in some embodiments, the second objective function satisfies the following condition:

[0148]

[0149] in, d(p) represents the cluster center of category h. i ,μ h ) represents sample p i The Euclidean distance between the cluster centers and h represents the category obtained through each preset clustering model, θ is the indicator function, and θ(y) is the distance between the cluster centers and h. i =h) is used to determine sample p i Category y i Is it h? If not, then it is 0. P is the first sample set. Characterizing the actual sample p i Category y i and sample p j Category y j The penalty parameter θ(y) for classifying the same items as different categories. i ≠y j Used to determine sample p i Category y i and sample p j Category y j If they are the same, the value is 0. Characterizing the actual sample p i Category y i and sample p j Category y j The penalty parameter θ(y) for classifying different items as belonging to the same category. i =y j Used to determine sample p i Category y i and sample p j Category y jWhether they are the same or different, the value is 0 if they are different. M and N identify the first sample set, and k represents the number of sample categories.

[0150] Each module of the model determination device provided in this application embodiment can realize the functions of each step of the model determination method provided above, and can achieve its corresponding technical effects. For the sake of brevity, it will not be described in detail here.

[0151] Based on the same inventive concept, embodiments of this application also provide an electronic device.

[0152] Figure 4 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.

[0153] An electronic device may include a processor 401 and a memory 402 storing computer program instructions.

[0154] Specifically, the processor 401 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0155] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory.

[0156] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this disclosure.

[0157] The processor 401 reads and executes computer program instructions stored in the memory 402 to implement any of the model determination methods in the above embodiments.

[0158] In one example, the electronic device may also include a communication interface 403 and a bus 410. For example, Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.

[0159] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0160] Bus 410 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Linear Predictive Coding (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (Peripheral Component Interconnect-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VESA Local Bus, VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application contemplates any suitable bus or interconnection. The electronic device can perform the model determination method in the embodiments of this invention, thereby implementing the model determination method described above.

[0161] Furthermore, in conjunction with the model determination methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the model determination methods in the above embodiments.

[0162] This application also provides a computer program product, wherein the instructions in the computer program product, when executed by a processor of an electronic device, cause the electronic device to perform various processes implementing the determination method embodiment of any of the above models.

[0163] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0164] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0165] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0166] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0167] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A method for determining a model, characterized in that, include: Obtain text data from multiple dimensions; The text data is converted into vector data to obtain the second sample set; Randomly select a second preset number and preset dimension of vectorized data from the second sample set to obtain a first sample set of a first preset number; Obtain a first sample set of a first preset number and a preset clustering model, wherein the first sample set includes vectorized data of a second preset number and preset dimensions; The first sample set is clustered using the preset clustering model to obtain clustering results, and the first objective function value of the preset clustering model is calculated. The first objective function value characterizes the degree of clustering of the clustering results. The first clustering model with the smallest first objective function value is added to the clustering model set. A second clustering model is selected from the preset clustering models and added to the clustering model set. If the number of cluster models in the cluster model set is less than a third preset number, a third cluster model is selected from the preset cluster models and added to the cluster model set until the number of cluster models in the cluster model set is not less than the third preset number, and all cluster models in the cluster model set are determined as target cluster models. Wherein, the second objective function values ​​of the first clustering model and the second clustering model are less than the first objective function value of the first clustering model, and the third objective function values ​​of the first clustering model, the second clustering model and the third clustering model are less than the second objective function value.

2. The method for determining the model according to claim 1, characterized in that, Also includes: Based on the clustering results of each clustering model in the target clustering model, the third preset number of adjacency matrices are generated; Calculate the weighted average of the adjacency matrices of the third preset number; Based on the weighted average value, the category of each sample in the first sample set of the first preset number is determined.

3. The method for determining the model according to claim 1, characterized in that, The preset clustering model includes a first sub-clustering model, and the method further includes: In the absence of labeling information in the first sample set, the first sample set is clustered using the first sub-clustering model to obtain the first clustering result, and the fourth objective function value of the first sub-clustering model is calculated using the first objective function; wherein, the labeling information includes information that the first sample and the second sample in the first sample set belong to the same category; The fourth clustering model with the smallest fourth objective function value is added to the first clustering model set. A fifth clustering model is selected from the first sub-clustering models and added to the first clustering model set. If the number of cluster models in the first cluster model set is less than the third preset number, a sixth cluster model is selected from the first sub-cluster models, and the sixth cluster model is added to the first cluster model aggregation until the number of cluster models in the first cluster model set is not less than the third preset number, and all cluster models in the first cluster model set are determined as the first target cluster model. Specifically, the fifth objective function values ​​of the fourth clustering model and the fifth clustering model are less than the fourth objective function value of the fourth clustering model, and the sixth objective function values ​​of the fourth clustering model, the fifth clustering model, and the sixth clustering model are less than the fifth objective function value.

4. The method for determining the model according to claim 1, characterized in that, The preset clustering model includes a second sub-clustering model, and the method further includes: When the first sample set includes labeling information, the first sample set is clustered using the second sub-clustering model to obtain a second clustering result, and the seventh objective function value of the second sub-clustering model is calculated using the second objective function; wherein, the labeling information includes information that the first sample and the second sample in the first sample set belong to the same category; The seventh clustering model with the smallest seventh objective function value is added to the second clustering model set. An eighth clustering model is selected from the second sub-clustering models and added to the second clustering model set. If the number of cluster models in the second cluster model set is less than the third preset number, a ninth cluster model is selected from the second sub-cluster models and added to the second cluster model set until the number of cluster models in the second cluster model set is not less than the third preset number, and all cluster models in the second cluster model set are determined as the second target cluster model. Specifically, the eighth objective function value of the seventh clustering model and the eighth clustering model is less than the seventh objective function value of the seventh clustering model, and the ninth objective function value of the seventh clustering model, the eighth clustering model, and the ninth clustering model is less than the eighth objective function value.

5. The method for determining the model according to claim 3, characterized in that, The first objective function satisfies the following condition: in, This represents the cluster center of category h. sample The Euclidean distance between the cluster centers and the cluster center, h represents the number of clusters obtained through each preset clustering model. For indicator functions, Used to determine samples Category Is it h? If not, it is 0. P is the first sample set, and k represents the number of sample categories.

6. The method for determining the model according to claim 4, characterized in that, The second objective function satisfies the following condition: in, This represents the cluster center of category h. sample The Euclidean distance between the cluster centers and the cluster center, h represents the number of clusters obtained through each preset clustering model. For indicator functions, Used to determine samples Category Is it h? If not, then it is 0. P is the first sample set. Characterizing actual samples Category and samples Category Penalty parameters for classifying the same data as different categories. Used to determine samples Category and samples Category If they are the same, the value is 0. Characterizing actual samples Category and samples Category The penalty parameter for classifying different categories as the same. Used to determine samples Category and samples Category Whether they are the same or different, the value is 0 if they are different. M and N identify the first sample set, and k represents the number of sample categories.

7. A model determining device, characterized in that, include: The acquisition module is used to acquire text data from multiple dimensions; The conversion module is used to convert the text data into vectorized data to obtain the second sample set; The selection module is used to randomly select a second preset number and a preset dimension of vectorized data from the second sample set to obtain a first sample set of a first preset number; The acquisition module is used to acquire a first sample set of a first preset number and a preset clustering model, wherein the first sample set includes vectorized data of a second preset number and a preset dimension; The determination module is used to cluster the first sample set using the preset clustering model to obtain clustering results, and to calculate the first objective function value of the preset clustering model, wherein the first objective function value characterizes the degree of clustering of the clustering results; An addition module is used to add the first clustering model with the smallest first objective function value to the clustering model set, select a second clustering model from the preset clustering models, and add the second clustering model to the clustering model set; The adding module is further configured to, when the number of cluster models in the cluster model set is less than a third preset number, select a third cluster model from the preset cluster models, add the third cluster model to the cluster model set, until the number of cluster models in the cluster model set is not less than the third preset number, and determine all cluster models in the cluster model set as target cluster models; Wherein, the second objective function values ​​of the first clustering model and the second clustering model are less than the first objective function value of the first clustering model, and the third objective function values ​​of the first clustering model, the second clustering model and the third clustering model are less than the second objective function value.

8. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the method for determining the model as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the method for determining the model as described in any one of claims 1-6.

10. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device is able to perform the model determination method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Data index determination method and device, equipment and storage medium

    CN115249098A

  • Cloud monitoring stream data detection method and device based on clustering analysis and storage medium

    CN115454779A