Model application method and system based on equivalence class division hierarchical dimension reduction
By using equivalence class partitioning and hierarchical dimensionality reduction, complex high-dimensional datasets are decomposed into hierarchical smaller datasets, enabling the training of small models that can run on ordinary hardware environments. This solves the problem of high cost of high-dimensional models and expands the application scope of artificial intelligence.
Patent Information
- Application Number
- CN202511031708.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, models trained on complex, high-dimensional datasets have a large number of parameters, requiring high-performance GPUs, which leads to high costs and makes them difficult to apply in general environments, thus limiting the scope of artificial intelligence applications.
By using equivalence class partitioning and hierarchical dimensionality reduction, the dataset is decomposed into smaller datasets with hierarchical structures, training small models that can run in general environments, and solving complex problems through ensemble execution.
This lowers the barrier to entry for model applications, enabling models with complex datasets to run on ordinary hardware, reducing costs and expanding the scope of artificial intelligence applications.
Smart Images

Figure CN120930715A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a model application method and system based on equivalence class partitioning hierarchical dimensionality reduction. Background Technology
[0002] In practical applications of artificial intelligence, complex problems often involve high-dimensional datasets. Models trained on these datasets typically have a large number of parameters, requiring significant computing resources and relying heavily on high-performance GPUs. This makes them unsuitable for general environments, resulting in high marginal costs that are unaffordable for many cost-sensitive small and medium-sized enterprises, thus limiting the scope of AI applications. Therefore, reducing model complexity and lowering the application threshold for complex datasets is a pressing issue that needs to be addressed. Summary of the Invention
[0003] In view of this, this paper proposes a model application method and system based on equivalence class partitioning and hierarchical dimensionality reduction. By performing equivalence class partitioning and hierarchical dimensionality reduction on the dataset corresponding to the target problem, the dataset is decomposed into a series of smaller datasets with hierarchical structures. A model is trained for each smaller dataset, resulting in a series of smaller models that can run in general environments. The original problem is solved by the integrated operation of multiple smaller models, thereby reducing the threshold for model application.
[0004] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0005] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:
[0006] Figure 1 This is a schematic diagram of the technical route of the model application method based on equivalence class partitioning hierarchical dimensionality reduction provided by the present invention.
[0007] Figure 2 This is a schematic diagram of the application method of the model based on equivalence class partitioning hierarchical dimensionality reduction provided by the present invention.
[0008] Figure 3 This is a schematic diagram of a hierarchical dataset provided according to the present invention.
[0009] Figure 4 This is a schematic diagram of a hierarchical dataset with labeled sub-datasets provided according to the present invention.
[0010] Figure 5This is a schematic diagram of the hierarchical model provided by the present invention.
[0011] Figures 6-11 These are schematic diagrams of the interface of a model application system provided according to two embodiments of the present invention.
[0012] Figure 12 This is a schematic diagram of the structure of a model application system based on equivalence class partitioning and hierarchical dimensionality reduction provided by the present invention. Detailed Implementation
[0013] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention. The present invention first provides a model application method 100 based on equivalence class partitioning hierarchical dimensionality reduction, such as... Figure 2 As shown, the application method 100 includes:
[0014] Step S110: Construct a hierarchical dataset. First, organize the background data corresponding to the target problem into a dataset, denoted as the full dataset S. Then, construct a hierarchical dataset based on the full dataset S using a hierarchical partitioning method. First, select the dimension with the lowest cardinality (greater than 1) among the dimensions of dataset S. Then, divide the dataset according to the different values of the dimension, with one value corresponding to one partition, thus dividing dataset S into several sub-datasets. Then, divide the sub-datasets in the same way until all dimensions have been partitioned. Finally, combine all the partitioned datasets according to the hierarchical relationship to construct a hierarchical dataset.
[0015] Specifically, step S110 can be implemented according to the following steps:
[0016] Step S111: Construct the complete dataset S. Conduct research on the target problem and collect background data. Based on the collected data, construct a sample dataset S = (s1, s2, ..., s...) containing n dimensions and m dimensions. n This is called the complete dataset, where the dimension set D = (d1, d2, ..., dn) m ), single sample s i =(v i1 ,v i2 ,…,v im ), s i It is a multiset, and elements may be repeated. The target model corresponding to the target problem is defined as the overall model M. The overall model M is constructed based on the entire dataset S, and the construction process will be explained step by step later.
[0017] Step S112: Discretize the entire dataset S. First, group each dimension of the entire dataset S into intervals. For dimension d... j value set d j =(v 1j ,v 2j ,…,v nj If d is a multiset, it may contain duplicate values. j If d is a set of continuous values, then d j Discretization is performed using methods such as the Likert scalar method; here, interval discrimination is used for discretization: [The text then continues with further details about the discretization process.] For set d j The minimum value in, For set d j The maximum value in d is d. j The upper and lower value ranges are Take c j As an interval number, the dimension d i The interval step size is:
[0018]
[0019] any interval in:
[0020]
[0021]
[0022] set d j =(v 1j ,v 2j ,…,v nj The set of intervals corresponding to each element in ) is:
[0023]
[0024] Using d′ j Replace d j This completes the discretization process for that dimension.
[0025] Construct D′=(d′1,d′2,…,d′) m When d i When ((i=(1,2,…,m)()) is a discrete value dimension, then d′ i =d i At this time, the set For d i The set after removing duplicate values, i.e., the general set with distinct elements; when d i When the dimension is continuous, then for d i Discretize the data into discrete value dimension d′i This completes the discretization of the entire dataset S.
[0026] Step S113: Perform hierarchical partitioning on the entire dataset S. First, select the dimension with the smallest cardinality (i.e., the dimension with the fewest discrete values) or the preferred partitioning dimension as the partitioning dimension. Then, divide the entire dataset S into corresponding subsets based on the number of discrete values in that dimension. The method for dimension-based partitioning is as follows:
[0027] Define d * For the dimension that needs to be divided, when d * When prioritizing dimensions, these are the special dimensions that should be prioritized for classification. These can be highly salient characteristics (such as material, shape, function, etc., dimensions that easily create type distinctions); when d * When the cardinality dimension is the minimum, it can be obtained through cardinality calculation. Let set P = (P1, P2, ..., P...). m ) element P i For the dataset S′, the corresponding dimension d i The set of discrete values S′ contained therein can be the entire dataset S or a subset of the dataset in the hierarchical structure. Its minimum cardinality dimension is calculated as follows:
[0028]
[0029] d * =(d i |P i =P * )
[0030] For the smallest cardinality dimension (or the preferred partitioning dimension) d * The equivalence classes of the contained numerical sets are constructed as follows:
[0031] Equivalence classes are constructed from equivalence relations, which can take many forms. One such equivalence relation is listed here. An equivalence relation can be defined as follows: In dataset S′, d j d is one dimension of S′ j The set of discrete values included For sample s′ i The corresponding dimension d in j element v′ ij and discrete values Where l = (1,2,…,n) j ),like but The equivalence class generated by the equivalence relation is denoted as . The set of all equivalence classes is called d. j Regarding the commercial collection of ~,
[0032] Finally, let the dimension d to be partitioned be... * =d j , through d j The dataset S′ is partitioned into equivalence classes based on the dimension d. j The set of discrete values contained therein Divided into in:
[0033]
[0034] Let S′ l Let l be a partition of S′, where l = (1, 2, ..., n) j ), S′ l The corresponding equivalence class, S′ l Number of rows contained Then we have:
[0035]
[0036] Because S′ l In d j All values of the dimension are Therefore d j Dimension in S′ l It is a single-valued dimension. Value at S′ l It is universal and lacks comparative and analytical significance; therefore, it is through S′ l Delete d j Dimensionality reduction is achieved. Definition To remove one dimension from the dimension set D of the dataset, the following processing is performed:
[0037]
[0038] The above completes the process through dimension d. j A subset of the lower equivalence class partition, i.e., through dimension d j Different values of d divide the dataset S′ into several subsets, each subset corresponding to a d. j The discrete values of the dimension are used to divide the data into subsets, and dimensionality reduction is achieved by deleting the column of that dimension.
[0039] Step S114, recursive partitioning. First, partition the resulting subset S′. lFollowing the method in step S113, find the next dimension for partitioning, then find the next dimension for partitioning the resulting subset, and so on recursively until all dimensions have been partitioned (deleted). Continue recursively partitioning the subset according to the remaining unpartitioned (undeleted) dimensions. Let the t-th subset of the r-th layer have u subsets, then the subset S... rt satisfy:
[0040]
[0041] d rk For a subset S rk The corresponding partitioning dimensions This represents adding a dimension from set D to dataset S, essentially restoring the dimensions removed during dimensionality reduction. This completes the process as follows: Figure 3 The hierarchical dataset construction is shown.
[0042] Step S120: Construct the hierarchical master model M. First, label the subsets of data that need to be trained. Then, assign a model to the labeled dataset, extract the parameter set of the assigned model, and train the sub-models using the "model training system". Finally, integrate the sub-models according to the subset structure to form a hierarchical model structure.
[0043] Specifically, step S120 can be implemented according to the following steps:
[0044] Step S121, Label the dataset. After constructing the hierarchical dataset, label the dataset by observing its structure, pruning smaller datasets and dividing larger datasets into several appropriately sized datasets, such as... Figure 4 As shown, the bold red subsets represent labeled datasets, while unlabeled datasets will not be used for training. In the tree structure, each branch has only one labeled subset (i.e., only one branch model is trained per branch). Through the above steps, the set of labeled datasets is obtained: b is the number of labeled datasets, n s b = n, where b is the number of all subsets of the dataset, and b = n is the number of subsets of the dataset when the hierarchical dataset has only one level. s .
[0045] Step S122, Sub-model creation. For each labeled subset of data, specify the model to be trained. This is generally a discriminative model, such as a multilayer perceptron or convolutional neural network, or it can be a combined model (i.e., a combination of multiple models), thus obtaining the model set. M s The elements in the middle correspond to the labeled datasets: Then, the parameter set of each model is refined (at this time, the model parameters still do not have parameter values and belong to the formal parameter state). For example, the parameter set of a multilayer perceptron is the number of layers, the number of neurons, the activation function, the solver, the regularization coefficient, the learning rate, etc. If it is a combination of multiple models, the parameter lists of these multiple models are merged into a set.
[0046] Step S123, Model Training. The model and parameter list constructed in step S122 are input into the "Model Training System" designed and developed in "Model Training Method and Model Training System Based on Reinforcement Learning (Application No.: 2025104266479)" to train the model and output the model parameter values, thus obtaining an entity model with parameter values.
[0047] The patent application "Model Training Method and System Based on Reinforcement Learning (Application No.: 2025104266479)" is currently in the "Invention Publication (Publication No.: CN120338032A)" stage, which expires on July 19, 2025. It discloses a model training method and system, proposing a reinforcement learning-based model training method that can batch-set parameter combinations for models, automating the model training and evaluation process, seeking the optimal parameter combination to output an optimized model. The technical content of this patent application is as follows:
[0048] The parameter set of the model to be trained is used to perform localized parameter combinations within the bounded range of each parameter to obtain an initial parameter combination set; a performance evaluation model oriented towards performance metrics is constructed based on the dataset and the initial parameter combination set; a reinforcement learning perceptual response model is constructed based on the performance evaluation model, and the parameters in the baseline parameter combination are randomly replaced iteratively to determine the optimal performance parameter combination that can produce the maximum reward value; the model to be trained is trained using the optimal performance parameter combination, and the performance evaluation model is iterated to obtain the optimized target model.
[0049] Meanwhile, the patent provides a model training system that enables the above-mentioned technical steps to be implemented through a computer system. This step uses this system to complete the model set M. s =(m1,m2,…m b ) training.
[0050] Step S124, construct the hierarchical model. Based on the hierarchical dataset described in step S121 ( Figure 4 Construct a hierarchical model set (one sub-dataset corresponds to one sub-model), and build models corresponding to the three types of sub-datasets:
[0051] I. Labeled Subsets: The model corresponding to the labeled subsets is an entity model (a trained model with parameter values), m′=(m i |Si ), i = (1, 2, ..., k b ).
[0052] II. Unlabeled parent subsets: Unlabeled parent datasets that are also labeled datasets. The corresponding model is a composite model (composed of entity models or other composite models). The dataset S is a dataset with u subsets in the r-th layer and t-th subset. rt The corresponding sub-model is:
[0053] m rt =(m (r+1)1 ,m (r+1)2 ,…,m (r+1)u )
[0054] III. Unlabeled Sub-datasets: Sub-datasets that are not labeled and are subordinate to labeled subsets. These subsets are pruned and do not correspond to any model (e.g., ...). Figure 5 (As shown).
[0055] Step S130: Apply the overall model M. First, the overall model M is converted into a computer program and integrated into the computer system. Then, the target problem is quantified into standardized data s′ to be analyzed. Next, the hierarchical structure of the overall model M is used to perform hierarchical matching on s′ until the entity model m′ is matched. s′ is input into m′ for calculation, and the result is the output of the overall model M. The result is returned graphically in the computer system.
[0056] Specifically, step S130 can be implemented according to the following steps:
[0057] Step S131, create standardized data. Quantify the target problem to be analyzed and collect data including the dimension set D = (d1, d2, ..., d...). m The data of each element in the dataset S are used to form a dataset s′=(v′1,v′2,…,v′) that is isomorphic to the elements in the entire dataset S. m ).
[0058] Step S132, hierarchical matching. Perform hierarchical matching on s′. For each hierarchical node s′ passes through, select the corresponding dimension value and choose the next hierarchical branch until a matching entity model is found for analysis and calculation, returning the result. From step S122, we know that each sub-model in the hierarchical structure corresponds to a subset of data. From step S113, we know that each subset of data corresponds to a partition of the parent dataset in the hierarchical structure based on the equivalence relation ~. This is because the equivalence class of the subset can be determined by finding the corresponding subset through the sub-models on the main model M. For data s′=(v′1,v′2,…,v′…) mIn the hierarchical structure of the main model M, processing is performed layer by layer. When s′ enters the first layer, the total dataset S corresponding to the main model M is divided by dimension d. k Where k = (1,2,…,m), and dimension d k The set of discrete values included The entire dataset S is divided into s′=(v′1,v′2,…,v′ m In dimension d k The value on is v′ k Thus, the subset of data s′ corresponding to this layer is obtained. Because of the one-to-one correspondence between the subset and the sub-model, the sub-model m′ can be determined from the subset p′, where m′∈M. s =(m1,m2,…m b If the subset p′ is a labeled subset, then m′ is the entity model trained on the subset. The result of the operation on s′ by the entity model m′ is the result of the main model M. If the subset p′ is not labeled, then m′ is the combined model. The process is repeated layer by layer. The value of s′ on the corresponding dimension of the current layer subset is obtained. The corresponding next layer subset is matched. The corresponding model is found through the subset. The model operation is performed according to whether the subset is a labeled subset or the process continues layer by layer. This process is repeated until a labeled subset and an entity model are matched to calculate and return the result.
[0059] Step S132: Output the results. After matching the final entity model m′ corresponding to the data s′ to be analyzed through the above steps, the result obtained by the calculation of m′ is returned to the main model M, and the main model M outputs the corresponding results through computer graphical interface or other means. Detailed Implementation
[0061] In addition, the technical roadmap of this invention can be referred to Figure 1 This demonstrates the overall logic of the present invention. In practical applications of the embodiments, the above-described solution of the present invention can be implemented according to the following steps:
[0062] Step S190: Full dataset (S111).
[0063] Organize the background data corresponding to the target problem into a dataset, denoted as the full dataset S, which is the content described in step S111.
[0064] Step S191: Hierarchical division (S112-3).
[0065] The entire dataset S is transformed into a hierarchical dataset using equivalence class partitioning. This is the content described in steps S112 and S113. In the diagram, "Sub-dataset 01", "Sub-dataset 02", "Sub-dataset 021", "Sub-dataset 022" (and "..." are descriptive terms, not referring to specific datasets, but rather to subsets within the hierarchical dataset. "Sub-dataset 01" and "Sub-dataset 02" are subsets of the entire dataset S, while "Sub-dataset 021" and "Sub-dataset 022" are subsets of "Sub-dataset 02". Based on this partitioning, the entire dataset S is transformed into a hierarchical structure.
[0066] Step S192: Label the dataset (S121)
[0067] The dataset to be used for model training is labeled, and the labeling results are as follows: Figure 4 As shown, the large dataset is divided into several smaller datasets by labeling the datasets, and the smaller datasets are also pruned, which is the content described in step S121.
[0068] Step S193: Construct the model (S122)
[0069] Specifying a formal model for the labeled dataset is what is described in step S122.
[0070] Step S194: Training (S123)
[0071] Model training using the "model training system" described in "Model Training Method and System Based on Reinforcement Learning" is the content described in step S123. The "output," "entity model a01," and "entity model a02" shown in the figure refer to the entity models trained by the aforementioned "model training system." "Entity model a01" and "entity model a02" are generic terms and do not represent specific models.
[0072] Step S195: Integration (S124)
[0073] Constructing a hierarchical model and integrating the entity models to create a combined model is the content described in step S124.
[0074] Step S196: Data to be analyzed (S131)
[0075] Constructing standardized data and applying it to the main model M is what is described in step S131.
[0076] Step S197: Hierarchical matching (S132)
[0077] By matching the data s′ to be analyzed layer by layer, when s′ passes through a level node, the dimension value corresponding to this level node is taken to select the level branch to enter, until the entity model is matched, the analysis operation is performed and the result is returned, which is the content described in step S132.
[0078] Step S198: Output the result (S133)
[0079] The result of the calculation is output, which is the content described in step S133.
[0080] Example
[0081] To further explain the present invention, two specific application embodiments are provided below. Embodiment 1 is an assessment of the development capability of a specific module in software development, specifically describing the implementation process and application logic of the present invention; Embodiment 2 is the shape inspection stage of quality inspection in industrial production, specifically describing the breakthrough of the present invention in lowering the threshold for model application of image detection technology.
[0082] Example 1
[0083] In the field of software development, different types and levels of software development work often place different demands on engineers. For example, low-level system development requires a more solid foundation in development skills, architecture-oriented system development requires a more comprehensive technical perspective, while business application-oriented development requires a detailed understanding of the business. Defining the specific capability requirements for developing and maintaining a system is a complex matter, and judging whether an engineer meets those requirements is even more difficult. Often, assessment can only be made based on the engineer's past work experience.
[0084] Against this backdrop, CR Company needs to assess whether engineers are qualified to develop and maintain the software product (PRS system). The PRS system comprises numerous software modules, many of which are also used in other systems. For example, the efp_efm_0262 module exists not only in the PRS system but also in the IMP and WSP systems. Software engineers typically do not have experience with all relevant modules of a system, but many modules share similarities and connections. Therefore, if an engineer's past development experience involves certain software modules, they can more easily understand other related and similar modules, which greatly aids in understanding the overall system development and makes them more capable of handling the maintenance work.
[0085] The objective of this embodiment is to determine a software engineer's suitability for PRS development by examining their existing development experience. Therefore, this embodiment examines the performance of all engineers who have participated in PRS system development, as well as their prior work experience before their initial participation in PRS system development. By correlating previous work experience with subsequent performance in PRS, suitable work experience for PRS system development is identified. This embodiment uses an MLP (Multilayer Perceptron) model applied through the method provided by this invention, ultimately outputting a competency evaluation result.
[0086] Step S310 (this step corresponds to step S111 described above), collect dataset S p .
[0087] The dataset was prepared by taking the work experience of software engineers who had previously worked on PRS development before taking on this role, as well as their final job evaluations (competence assessment). Figure 6 The table columns efp_efm_0xxx and efp_etp_0xxx (where 0xxx is a data code) in the figure represent module numbers. Table content 1 indicates that the engineer had complete understanding of the module before working on PRS system development, 0 indicates no relevant experience, and values between 0 and 1 represent the degree of understanding. The column "Competent or Incompetent" represents the engineer's post-development performance evaluation, with 1 indicating competence and 0 indicating incompetence, and values between 0 and 1 representing the degree of competence. This dataset records engineers' work experience before working on PRS system development and their post-development performance evaluation. Therefore, it can be used to create a model for "predicting whether an engineer with known work experience is competent in developing a PRS system."
[0088] Step S311 (this step corresponds to steps S112, S113, and S121 mentioned above), hierarchical division and dataset labeling.
[0089] The dataset S collected in step S310 p Perform hierarchical partitioning, with dimension set D = (d1, d2, ..., d...). m That is, dataset S pThe columns efp_efm_0xxx and efp_etp_0xxx (0xxx being data symbols) are discretized as described above. (In this embodiment, many columns contain only 0 and 1 values, which are already discrete columns and therefore do not require further discretization. Some columns contain values between 0 and 1, requiring interval discretization.) The column with the smallest cardinality is selected (since many columns in this embodiment only contain 0 and 1, the cardinality is 2; therefore, a suitable column with the smallest cardinality is selected manually based on its importance, i.e., its weight in influencing the result). After selecting the column with the smallest cardinality, the data is partitioned. This process is repeated until all columns have been partitioned (or the loop can be stopped manually, and only some important columns can be partitioned). Next, the constructed hierarchical dataset is labeled (in this example, the lower-level datasets of the labeled dataset are deleted). The final result is as follows... Figure 7 As shown, the datasets “efp_etp_0266(=(1)” and “efp_etp_0266(=(0)” are subsets of the dataset “efp_efm_0262(=(1)”. Here, “1” and “0” are two discrete values of the dimension “efp_etp_0266”, which constitute two equivalence classes of “efp_etp_0266” under “efp_efm_0262(=(1)”, thus partitioning the dataset.
[0090] Step S312 (this step corresponds to step S122 mentioned above), model training.
[0091] For each labeled subset of the dataset, an MLP (Multilayer Perceptron) model is built, the model parameters are extracted, and the model is trained using the method described in step S122. Figure 8 (This is only an explanation of the internal operation of the model training method described in step S122, "Model Training Method and Model Training System Based on Reinforcement Learning (Application No.: 2025104266479)," and is not the model trained on the data in this example.) This is a diagram illustrating the model training results for one subset of the dataset. The columns in the diagram—number of layers, number of neurons, activation function, solver, regularization coefficient, learning rate, initial learning rate, maximum number of iterations, shuffling order, random seed, and convergence threshold—are all parameters of the MLP model. The optimal model is selected from the models trained in batch training as the training result for the labeled subset of the dataset.
[0092] Step S313 (corresponding to step S124 described above) involves model ensemble construction of a hierarchical model tree. Through step S312, the models trained on each labeled subset of the dataset are ensembled into the unlabeled dataset at the higher level, thus forming a hierarchical model tree. As indicated by the MLP model specified in step S312, this hierarchical model tree is an ensemble application of MLP models.
[0093] Step S314 (this step corresponds to step S130 mentioned above), model application.
[0094] First, data on the work experience of engineers is collected according to the dataset's data standards (dimension number correspondence), corresponding to the process described in step S131. Then, this work experience data is input into the hierarchical model constructed in step S312, which has a hierarchical structure... Figure 7 The hierarchical datasets are consistent (both share the equivalence class partitioning method). If the engineer has work experience with module efp_efm_02581, then the partition "efp_efn_02581=1" is entered. Then it is determined whether the engineer has work experience with efp_efm_0262. If so, the partition "efp_efm_0262=1" is entered. Finally, the partition with the marked name (such as "efp_efm_0260(=1" or "efp_efm_0260(=0")) is matched to find the corresponding entity model for calculation and output of results.
[0095] Example 2
[0096] In the industrial manufacturing sector, it is frequently necessary to perform visual inspections on work-in-process to promptly identify products with quality defects and rework or discard them to prevent the formation of non-conforming products. Traditional visual inspection relies heavily on manual labor, which is not only costly but also prone to missed or incorrect inspections. Against this backdrop, QS Company decided to use computer image algorithms as an auxiliary tool for visual inspection. However, the company's products have a wide sales volume, large output, and many models, specifications, and materials. If traditional image algorithms were used, inspecting a large number of products together would require a model with massive parameters for recognition and inspection, and running such a large model would pose a significant challenge to the hardware and software environment. Therefore, QS Company adopted the method described in this invention to perform visual inspection on gears on the production line. By using an equivalence class-based hierarchical dimensionality reduction method, the dataset is segmented and reduced in size. A lightweight CNN (Convolutional Neural Network) algorithm is used as the entity model, enabling visual inspection of the gears produced by the company to be performed in a typical hardware environment.
[0097] Step S320 (this step corresponds to step S111 described above), collect dataset S g By setting up a darkroom and high-speed image acquisition equipment on the production line of the appearance inspection process, appearance image data of the products to be inspected is collected. Simultaneously, by interfacing with the MES (Manufacturing Execution System), the production batch number and drawing information of the products to be inspected are obtained, and appearance-related parameters are collected to form a dataset S. g (like Figure 10 (As shown). Figure 10 The dataset S shown gIn the image, the "Qualified or Not" column indicates whether the product represented by the image is qualified; the "Top View" column contains the image data of the product to be inspected; the columns "Tooth Line (including values for spur gears, helical gears, herringbone gears, and curved gears)," "Tooth Profile (including values for involute, cycloid, and circular arc)," "Working Conditions (including values for cylindrical, bevel gears, worm gears, and racks)," "Transmission Method (including values for parallel shafts, intersecting shafts, and staggered shafts)," "Structure (including values for planetary, modified, and drum-shaped)," and "Material (including M16 and M22)" are all discrete columns; while "Module," "Number of Teeth," "Pitch Circle Diameter," "Addendum Circle Diameter," "Root Circle Diameter," and "Tooth Width" are all continuous value columns.
[0098] Step S321 (this step corresponds to steps S112, S113, and S121 mentioned above): hierarchical partitioning and dataset labeling. The dataset S collected in step S320... g Divide into hierarchical levels, such as Figure 11 As shown, equivalence classes are defined by "material", "structure", "tooth profile", "transmission method" and "working conditions" respectively, and the subset of datasets divided by the equivalence classes under "working conditions" are used as the subset of datasets for training the entity model.
[0099] Step S322 (corresponding to steps S124 and S130 described above): Model training and application. A CNN (Convolutional Neural Network) model is built for each labeled subset of the dataset. Model parameters are extracted, and the image data is used as input data for CNN model training, as described in step S122. The model is then applied using the method described in S130. During application, the image data is input into the corresponding hierarchical model after hierarchical matching, and the results are output.
[0100] It should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. The above are merely embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
[0101] On the other hand, the present invention also provides a "model application system (900) based on equivalence class partitioning hierarchical dimensionality reduction", which can be implemented based on the "model application method 100 based on equivalence class partitioning hierarchical dimensionality reduction" described in the present invention. Figure 12 As shown, the system (900) includes:
[0102] Device 910, a hierarchical dataset management device. Based on the method described in step S110 of this invention, it is used to construct and manage hierarchical datasets, such as... Figure 6 and Figure 10 For dataset data list, Figure 7 and Figure 11 This is a list of hierarchical dataset structures; the interfaces shown are all application interfaces of this system.
[0103] Device 920 is a hierarchical model training and integration management device. Based on the method described in step S120 of this invention, it is used for model training on a labeled dataset, integrating the system described in "Model Training Method and System Based on Reinforcement Learning (Application No.: 2025104266479)" for model training, and simultaneously integrating the trained entity models to construct a hierarchical model tree, such as... Figure 8 and Figure 9 A diagram illustrating the model training process;
[0104] Device 930 is a hierarchical model application and output presentation device. Based on the method described in step S130 of this invention, the model is applied and the result of the model calculation is output and presented.
Claims
1. A model application method based on equivalence class partitioning and hierarchical dimensionality reduction, characterized in that, The application method includes: Establish the complete dataset S = (s1, s2, ..., s n Using equivalence class partitioning, hierarchical dimensionality reduction is performed on the entire dataset S to construct a hierarchical dataset; In the hierarchical dataset, a subset of datasets is labeled and a model is assigned to it. The assigned model is trained using a reinforcement learning-based model training method. The trained entity models are then hierarchically integrated to construct a hierarchical model tree containing the combined model and the entity model, denoted as the total model M. The overall model M uses a hierarchical matching method to match the input data to be analyzed with entity models, performs calculations on the matched entity models and returns the results. The above process is implemented through a computer system and the calculation results are output in a graphical interface.
2. The model application method according to claim 1, characterized in that, The process of constructing a hierarchical dataset by performing hierarchical dimensionality reduction on the entire dataset S includes: Define D = (d1, d2, ..., d m Let S be the set of dimensions of the entire dataset S. First, all non-discrete dimensions of the entire dataset S are discretized. Then, the smallest cardinality dimension or the preferred partitioning dimension is taken as the partitioning dimension to perform a recursive hierarchical dimensionality reduction partitioning of the entire dataset S. That is, for the sub-datasets partitioned based on equivalence classes, the next smallest cardinality dimension or the preferred partitioning dimension is taken as the partitioning dimension for recursive partitioning. The dimension is then deleted from the partitioned sub-datasets to complete the dimensionality reduction process until all dimensions have been partitioned or the recursion stops by judgment.
3. The model application method according to claim 1, characterized in that, The step of constructing a hierarchical model tree containing combined models and entity models by hierarchical ensemble of the trained entity models includes: The entity model trained on the labeled subset is the leaf node of the hierarchical model tree, and the entity model is the model trained on the dataset. On the hierarchical model tree, the lower-level dataset nodes corresponding to the labeled subset are pruned. A combined model is built on the node corresponding to the upper-level dataset of the labeled subset. The combined model is a model formed by integrating one or more entity models and other combined models. Here, the combined model is integrated by the models corresponding to its child nodes on the hierarchical model tree.
4. The model application method according to claim 1, characterized in that, The process of matching the input data to be analyzed with an entity model includes: Define the data to be analyzed as s′=(v′1,v′2,…,v′) m When s′ is input into the main model M, s′ will enter the corresponding sub-models on the branch nodes of the main model M layer by layer, including the combined model and the entity model. When s′ passes through a level node, when it matches the combined model, it takes the sub-data set divided by the dimension value corresponding to this level node and enters the level branch of the sub-data set again, until it matches the entity model and stops. The entity model is used to analyze and calculate s′ and return the calculation result.
5. The model application method according to claim 2, characterized in that, The step of using the minimum cardinality dimension or the preferred partitioning dimension as the partitioning dimension to perform recursive hierarchical dimensionality reduction partitioning of the entire dataset S includes: Define d * For the dimensions that need to be divided, when D contains special dimensions that represent important significance, then d * It can be one of these important dimensions, becoming the preferred dividing dimension; otherwise, d * It can be the smallest cardinal dimension obtained through cardinality calculation, let set P = (P1, P2, ..., P...). m ) element P i For the dataset S′, the corresponding dimension d i The set of discrete values S′ contained therein can be the entire dataset S or a subset of the dataset in the hierarchical structure. Its minimum cardinality dimension is calculated as follows: d * =(d i |P i =P * ) When d * =d j At that time, through d j The dataset S′ is partitioned into equivalence classes based on the dimension d. j The set of discrete values contained therein Divided into in: Define S′ l Let l be a partition of S′, where l = (1, 2, ..., n) j ), S′ l The corresponding equivalence class, S′ l Number of rows contained have: definition To remove one dimension from the dimension set D of the dataset, the following processing is performed: This completes the dimensionality reduction.
6. A model application system based on equivalence class partitioning and hierarchical dimensionality reduction, characterized in that, The model application system includes: Device 910, a hierarchical dataset management device, is used to build and manage hierarchical datasets; Device 920, a hierarchical model training and integration management device, is used for model training of labeled datasets, integrating the system described in "Model Training Method and Model Training System Based on Reinforcement Learning (Application No.: 2025104266479)" for model training, and simultaneously integrating the trained entity models and constructing a hierarchical model tree; Device 930 is a hierarchical model application and output presentation device that outputs and presents the results of model calculations after applying the model.
Citation Information
Patent Citations
Model training method and model training system based on reinforcement learning
CN120338032A