Hybrid Expert Multi-Class Classification Method and System Combining Embedded Dual-Model Organizational Structure

By combining the hybrid expert multi-classification method embedded in the dual-model organizational structure, using the data distribution consistency threshold and the hybrid expert mechanism to construct a scientific value sentence recognition and multi-classification model, the problems of low efficiency and insufficient accuracy of scientific literature value sentence classification in the existing technology are solved, and efficient and accurate multi-dimensional analysis is achieved.

CN119646224BActive Publication Date: 2025-07-22DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411805132.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-07-22
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

The existing scientific literature value sentence classification methods have problems such as low processing efficiency and difficulty in ensuring accuracy when processing scientific literature.

Method used

By combining the hybrid expert multi-classification method embedded in the dual-model organizational structure, the pre-trained model is deeply learned by using data distribution consistency threshold and scientific value sentence training corpus, a scientific value sentence recognition model is constructed, and a scientific value sentence multi-classification model is built based on the hybrid expert mechanism, including academic value, application value and innovative value expert sub-model and gated network, encapsulated as a scientific value sentence detection dual-model.

Benefits of technology

It realizes efficient and accurate classification of scientific value sentences, and improves the analysis accuracy of multi-dimensional academic, application and innovative value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646224B_ABST
    Figure CN119646224B_ABST
Patent Text Reader

Abstract

The present invention discloses a hybrid expert multi-classification method and system combining an embedded dual-model architecture, which is applied to the field of data processing technology. By performing deep learning on a pre-trained model based on a data distribution consistency threshold and a corpus of scientific value sentence training materials, a scientific value sentence recognition model is constructed. Based on a hybrid expert mechanism, a scientific value sentence multi-classification model is built, wherein the scientific value sentence multi-classification model includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network. The scientific value sentence recognition model and the scientific value sentence multi-classification model are encapsulated into a scientific value sentence detection dual model to obtain a scientific and technological literature. The scientific and technological literature is input into the scientific value sentence detection dual model to obtain a comprehensive classification result of scientific value sentences. It solves the technical problems of low processing efficiency and difficult accuracy guarantee in the existing scientific literature value sentence classification method when dealing with scientific and technological literature in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a hybrid expert multi-classification method and system combining an embedded dual-model architecture. Background Art

[0002] With the continuous increase of scientific and technological literature, how to efficiently and accurately identify and classify scientific value sentences in the literature, such as academic value, application value, and innovation value sentences, has become an important issue in scientific research and literature management. Traditional literature classification methods mostly rely on manual annotation and rule setting. Although they can partially meet the requirements, they are often inefficient and inaccurate in processing large-scale literature.

[0003] Therefore, in the prior art, the existing scientific literature value sentence classification methods have technical problems of low processing efficiency and difficult accuracy guarantee when processing scientific and technological literature. Summary of the Invention

[0004] The present application provides a hybrid expert multi-classification method and system combining an embedded dual-model architecture, which solves the technical problems of low processing efficiency and difficult accuracy guarantee in the existing scientific literature value sentence classification methods when processing scientific and technological literature. By introducing a hybrid expert mechanism and optimizing data distribution consistency, the classification accuracy of scientific value sentences can be effectively improved, and the technical effect of efficiently and accurately analyzing multi-dimensional academic, application, and innovation values can be achieved.

[0005] The present application provides a hybrid expert multi-classification method combining an embedded dual-model architecture. The method includes: performing deep learning on a pre-trained model based on a data distribution consistency threshold and a training corpus of scientific value sentences to construct a scientific value sentence recognition model. Based on the hybrid expert mechanism, building a scientific value sentence multi-classification model, where the scientific value sentence multi-classification model includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network. Encapsulating the scientific value sentence recognition model and the scientific value sentence multi-classification model into a scientific value sentence detection dual-model. Obtaining a scientific and technological literature. Inputting the scientific and technological literature into the scientific value sentence detection dual-model to obtain a comprehensive classification result of scientific value sentences.

[0006] In an implementation manner, based on the data distribution consistency threshold and the training corpus of scientific value sentences, deep learning is performed on a pre-trained model to construct a scientific value sentence recognition model, including: The training corpus of scientific value sentences includes a scientific value sentence sample set and a non-scientific value sentence sample set. The training corpus of scientific value sentences is divided according to a predetermined ratio to obtain a first corpus training set, a first corpus test set, and a first corpus validation set. Based on the data distribution consistency threshold, data distribution optimization is performed on the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set. Based on the second corpus training set, the second corpus test set, and the second corpus validation set, the pre-trained model is trained, tested, and validated to generate the scientific value sentence recognition model.

[0007] In an implementation manner, based on the data distribution consistency threshold, data distribution optimization is performed on the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set, including: Calculating the positive and negative sample distributions of the first corpus training set, the first corpus test set, and the first corpus validation set respectively to obtain a first data distribution coefficient, a second data distribution coefficient, and a third data distribution coefficient. Based on the first data distribution coefficient, the second data distribution coefficient, and the third data distribution coefficient, a data distribution consistency evaluation is performed on the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a data distribution consistency coefficient. Determining whether the data distribution consistency coefficient is greater than the data distribution consistency threshold. If the data distribution consistency coefficient is greater than the data distribution consistency threshold, data distribution adjustment is performed on the first corpus training set, the first corpus test set, and the first corpus validation set according to the data distribution consistency threshold to generate the second corpus training set, the second corpus test set, and the second corpus validation set.

[0008] In an implementation manner, based on the mixture of experts mechanism, a multi-classification model for scientific value sentences is built, including: obtaining multi-classification metrics for scientific value sentences, where the multi-classification metrics for scientific value sentences include academic value, application value, and innovation value. Based on the multi-classification metrics for scientific value sentences, a corpus for scoring academic value sentences, a corpus for scoring application value sentences, and a corpus for scoring innovation value sentences are loaded. Based on the corpus for scoring academic value sentences, the academic value expert sub-model is trained. Based on the corpus for scoring application value sentences, the application value expert sub-model is trained. Based on the innovation value expert sub-model, the innovation value expert sub-model is trained. Based on the gating loss function, according to the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model, the gating network is constructed. Using the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model as multi-classification parallel nodes. Connecting the multi-classification parallel nodes and the gating network to generate the multi-classification model for scientific value sentences.

[0009] In an implementation manner, based on the gating loss function, according to the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model, the gating network is constructed, including: collecting output samples of the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain an expert output sample set. Collecting discriminant probability parameters corresponding to the expert output sample set to obtain a discriminant probability distribution. Using minimizing the gating loss as the training objective of the gating network. Based on the gating loss function and the training objective of the gating network, unsupervised training is performed on the expert output sample set and the discriminant probability distribution to generate the gating network.

[0010] In an implementation manner, the gating loss function is:

[0011]

[0012] where represents the gating loss function, represents the unsupervised loss, λ represents a predetermined balance weight, represents the minimization regularization term.

[0013] In the implementation manner, the scientific and technological literature is input into the dual model for detecting scientific value sentences to obtain a comprehensive classification result of scientific value sentences, including: locating the enrichment area of scientific value sentences based on the scientific and technological literature to obtain the enrichment area of value sentences. The enrichment area of value sentences is input into the recognition model of scientific value sentences to obtain the recognition result of scientific value sentences. The recognition result of scientific value sentences is input into the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain multi-type value discrimination results. The multi-type value discrimination results are input into the gating network to generate the comprehensive classification result of scientific value sentences.

[0014] The present application also provides a hybrid expert multi-classification system combined with an embedded dual model architecture, characterized in that the system includes:

[0015] A construction module for the recognition model of scientific value sentences, which is used to perform deep learning on a pre-trained model based on a data distribution consistency threshold and a training corpus of scientific value sentences to construct a recognition model of scientific value sentences.

[0016] A construction module for the multi-classification model of scientific value sentences, which is used to build a multi-classification model of scientific value sentences based on a hybrid expert mechanism, wherein the multi-classification model of scientific value sentences includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network.

[0017] A model encapsulation module, which is used to encapsulate the recognition model of scientific value sentences and the multi-classification model of scientific value sentences into a dual model for detecting scientific value sentences.

[0018] A module for obtaining scientific and technological literature, which is used to obtain scientific and technological literature.

[0019] A module for obtaining the comprehensive classification result, which is used to input the scientific and technological literature into the dual model for detecting scientific value sentences to obtain a comprehensive classification result of scientific value sentences.

[0020] A hybrid expert multi-classification method and system with an integrated dual-model architecture proposed in this application deep-learns a pre-trained model through a data distribution consistency threshold and a scientific value sentence training corpus to construct a scientific value sentence recognition model. Based on the hybrid expert mechanism, a scientific value sentence multi-classification model is built, where the scientific value sentence multi-classification model includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network. The scientific value sentence recognition model and the scientific value sentence multi-classification model are encapsulated into a scientific value sentence detection dual-model. A scientific and technological literature is obtained. The scientific and technological literature is input into the scientific value sentence detection dual-model to obtain a comprehensive classification result of the scientific value sentence. This solves the technical problems in the prior art that the existing scientific literature value sentence classification methods have low processing efficiency and difficult accuracy guarantee when processing scientific and technological literature. By introducing the hybrid expert mechanism and data distribution consistency optimization, the classification accuracy of scientific value sentences can be effectively improved, achieving the technical effect of efficient and accurate analysis of multi-dimensional academic, application, and innovation values. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments of the present disclosure will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0022] Figure 1 It is a schematic flowchart of the hybrid expert multi-classification method with an integrated dual-model architecture of the present invention;

[0023] Figure 2 It is a schematic structural diagram of the hybrid expert multi-classification system with an integrated dual-model architecture provided by an embodiment of the present application;

[0024] Description of the reference numerals: Scientific value sentence recognition model construction module 11, scientific value sentence multi-classification model construction module 12, model encapsulation module 13, scientific and technological literature acquisition module 14, comprehensive classification result acquisition module 15. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the following specifically gives the detailed description of the present application.

[0026] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0027] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first" and "second" are only used to distinguish similar objects and do not represent a specific order for the objects. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application.

[0028] Embodiments of this application provide a hybrid expert multi-classification method and system combined with an embedded dual-model architecture, as Figure 1 shown, the method includes:

[0029] Perform deep learning on a pre-trained model based on a data distribution consistency threshold and a corpus of scientific value sentence training data to construct a scientific value sentence recognition model.

[0030] Based on a hybrid expert mechanism, build a scientific value sentence multi-classification model, where the scientific value sentence multi-classification model includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network.

[0031] In scientific and technological literature, knowledge is not evenly distributed but exhibits a certain degree of concentration and regularity. In order to accurately classify academic value sentences, application value sentences, and innovation value sentences in scientific and technological literature, a deep learning is performed on a pre-trained model based on a data distribution consistency threshold and a training corpus of scientific value sentences to construct a scientific value sentence recognition model. Here, the data distribution consistency threshold is a preset value used to evaluate whether the sample distributions of the training, test, and validation data sets are balanced. When the class distribution of a data set is inconsistent with the class distribution of the overall data set, the data set is adjusted to meet this threshold. The training corpus of scientific value sentences is a text collection specially prepared for training the scientific value sentence recognition model, including sentence samples labeled as having or not having scientific value. Subsequently, based on the mixture of experts mechanism, a multi-classification model for scientific value sentences is built. Here, the multi-classification model for scientific value sentences includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network. The mixture of experts mechanism is an ensemble learning framework in which multiple sub-models (experts) evaluate and classify different categories of scientific value sentences according to their expertise. By combining multiple expert models and learning how to adaptively allocate to different experts according to the input, an adaptive decomposition and modeling of the task are achieved. Through the collaboration of multiple expert models, representative learning and decision output of the input are performed from different perspectives.

[0032] The method provided in the embodiment of this application further includes: The training corpus of scientific value sentences includes a scientific value sentence sample set and a non-scientific value sentence sample set. The training corpus of scientific value sentences is divided according to a predetermined ratio to obtain a first corpus training set, a first corpus test set, and a first corpus validation set. Based on the data distribution consistency threshold, data distribution optimization is performed on the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set. Based on the second corpus training set, the second corpus test set, and the second corpus validation set, the pre-trained model is trained, tested, and validated to generate the scientific value sentence recognition model.

[0033] The scientific value sentence training corpus is a set of example sentences selected from scientific and technological literature, which are marked as sentences with or without scientific value, namely the scientific value sentence sample set and the non-scientific value sentence sample set. Subsequently, the scientific value sentence training corpus is divided according to a predetermined ratio to obtain a first corpus training set, a first corpus test set, and a first corpus validation set, which are used for training, testing, and validating the model respectively. Among them, the predetermined ratio is a preset data division ratio, which can be divided according to the ratio of 70% training, 15% testing, and 15% validation. Further, based on the data distribution consistency threshold, the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set is optimized to ensure that each category is evenly distributed in the training, test, and validation sets, and a second corpus training set, a second corpus test set, and a second corpus validation set are obtained. Based on the second corpus training set, the second corpus test set, and the second corpus validation set, the pre-trained model is trained, tested, and validated. The pre-trained model is an untrained neural network model. Until the model is validated using the second corpus validation set and the output meets the preset accuracy rate, the validation is passed, and a scientific value sentence recognition model is obtained. The scientific value sentence recognition model is mainly used to judge whether the type of the input sentence is a scientific value sentence. The main function of this layer of the model is to filter out a large number of irrelevant sentences and provide high-quality corpus for subsequent classification tasks.

[0034] Calculate the positive and negative sample distributions of the first corpus training set, the first corpus test set, and the first corpus validation set respectively to obtain a first data distribution coefficient, a second data distribution coefficient, and a third data distribution coefficient. Based on the first data distribution coefficient, the second data distribution coefficient, and the third data distribution coefficient, evaluate the data distribution consistency of the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a data distribution consistency coefficient. Determine whether the data distribution consistency coefficient is greater than the data distribution consistency threshold. If the data distribution consistency coefficient is greater than the data distribution consistency threshold, adjust the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set according to the data distribution consistency threshold to generate the second corpus training set, the second corpus test set, and the second corpus validation set.

[0035] Based on the data distribution consistency threshold, perform data distribution optimization on the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set, including: calculating the positive and negative sample distributions of the first corpus training set, the first corpus test set, and the first corpus validation set respectively to obtain a first data distribution coefficient, a second data distribution coefficient, and a third data distribution coefficient. Among them, the first data distribution coefficient = the sample size of scientific value sentences in the first corpus training set ÷ the sample size of non-scientific value sentences in the first corpus training set, and the calculation processes of the second data distribution coefficient and the third data distribution coefficient are the same as that of the first data distribution coefficient. Based on the first data distribution coefficient, the second data distribution coefficient, and the third data distribution coefficient, perform a data distribution consistency evaluation on the first corpus training set, the first corpus test set, and the first corpus validation set. When performing the data distribution consistency evaluation, calculate the differences between the first data distribution coefficient and the second data distribution coefficient and the third data distribution coefficient respectively to obtain a data distribution consistency coefficient. Subsequently, determine whether the data distribution consistency coefficient is greater than the data distribution consistency threshold. The data distribution consistency threshold is the maximum judgment threshold for the preset data distribution difference consistency. When it is less than or equal to this threshold, the data distribution consistency of the corresponding two groups of data is relatively high. When it is greater than this threshold, the data distribution consistency of the corresponding two groups of data is relatively low, and data distribution optimization is required. If the data distribution consistency coefficient is greater than the data distribution consistency threshold, according to the data distribution consistency threshold, perform a distribution compensation adjustment on the first corpus training set, the first corpus test set, and the first corpus validation set according to the data distribution consistency coefficient, and supplement the existing difference ratio with data until it is less than the data distribution consistency threshold. Generate the second corpus training set, the second corpus test set, and the second corpus validation set.

[0036] The method provided by the embodiments of the present application further includes: obtaining multi-classification metrics for scientific value sentences, where the multi-classification metrics for scientific value sentences include academic value, application value, and innovation value. Based on the multi-classification metrics for scientific value sentences, loading the scoring corpus for academic value sentences, the scoring corpus for application value sentences, and the scoring corpus for innovation value sentences. Based on the scoring corpus for academic value sentences, training the academic value expert sub-model. Based on the scoring corpus for application value sentences, training the application value expert sub-model. Based on the scoring corpus for innovation value sentences, training the innovation value expert sub-model. Based on the gating loss function, constructing the gating network according to the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model. Using the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model as multi-classification parallel nodes. Connecting the multi-classification parallel nodes and the gating network to generate the multi-classification model for scientific value sentences.

[0037] Based on the mixture of experts mechanism, building a multi-classification model for scientific value sentences, including: obtaining multi-classification metrics for scientific value sentences, where the multi-classification metrics for scientific value sentences include academic value, application value, and innovation value. The academic value reflects the contribution of the sentence to the academic community, such as the importance of the proposed theory or experimental results. The application value reflects the potential or practical benefits of the research or discovery described in the sentence in practical applications. The innovation value is to evaluate the degree of innovation and originality of the research, method, or result mentioned in the sentence.

[0038] Subsequently, based on the multi-classification metrics of the scientific value sentences, the academic value sentence scoring corpus, the application value sentence scoring corpus, and the innovation value sentence scoring corpus are loaded. The academic value sentence scoring corpus: contains sentences marked as having high academic value and corresponding value scoring identifiers, such as sentences that deeply explore specific scientific issues. The application value sentence scoring corpus: contains sentences marked as having high application value and corresponding value scoring identifiers, such as sentences that describe research results with practical application prospects. The innovation value sentence scoring corpus: contains sentences marked as having high innovation value and corresponding value scoring identifiers, such as sentences that propose new theories or new methods. Further, based on the academic value sentence scoring corpus, the academic value expert sub-model is trained. Based on the application value sentence scoring corpus, the application value expert sub-model is trained. Based on the innovation value expert sub-model, the innovation value expert sub-model is trained. The academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model are all obtained through supervised training of a neural network model. Based on the gating loss function, the gating network is constructed according to the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model. Using the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model as multi-classification parallel nodes. Connecting the multi-classification parallel nodes and the gating network, the scientific value sentence multi-classification model is generated. Given the input x, the output y of the scientific value sentence multi-classification model can be expressed as:

[0039]

[0040] where f j (x) represents the output of the j-th expert model for the input x. g(x) j is the corresponding gating function. The gating mechanism plays the role of soft routing, adaptively allocating to different experts according to the input features, and combining their outputs with weights. g(x) is usually implemented using the Softmax function:

[0041]

[0042] where e j represents the gating logit value of the j-th expert, which can be obtained by linearly transforming x or mapping it through a feed-forward network. In the scientific value multi-classification task, this project has set up three expert models for academic value, application value, and innovation value respectively.

[0043] The method provided by the embodiment of the present application further includes: collecting output samples of the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain an expert output sample set. Collecting discriminant probability parameters corresponding to the expert output sample set to obtain a discriminant probability distribution. Taking minimizing the gating loss as the training objective of the gating network. Based on the gating loss function and the gating network training objective, performing unsupervised training on the expert output sample set and the discriminant probability distribution to generate the gating network.

[0044] Based on the gating loss function, constructing the gating network according to the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model, including: collecting output samples of the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain an expert output sample set. The expert output sample set is a set of output samples obtained from each expert sub-model (academic value, application value, innovation value expert sub-model). Each sample set contains scores or classification results for specific scientific value sentences, and these results reflect the attributes of the sentences in each value dimension. Collecting discriminant probability parameters corresponding to the expert output sample set to obtain a discriminant probability distribution. Taking minimizing the gating loss as the training objective of the gating network, and based on the gating loss function and the gating network training objective, performing unsupervised training on the expert output sample set and the discriminant probability distribution to generate the gating network.

[0045] The gating network plays a key routing role in the multi-classification model of scientific value sentences. Its goal is to adaptively assign different experts according to the characteristics of the value sentences. Different from the supervised training of the expert model, it is difficult to directly obtain a supervision signal for the gating network because the true expert assignment ratio for each value is unknown. Therefore, an unsupervised co-training method is adopted, using the output of the expert model as a soft label to guide the gating network to learn. In each training batch, we first use three expert models to perform inference on the input value sentences x1, x2,..., x b respectively, and obtain their discriminant probabilities p ij on their respective value dimensions. p ij = Sigmoid(f j (x i ), i = 1, 2,..., b, j = 1, 2, 3.

[0046] Then, taking these probability outputs as soft labels and performing JS divergence matching with the output of the gating network to obtain the unsupervised loss function of the gating network

[0047]

[0048] Among them, the definition of JS divergence is as follows:

[0049]

[0050] By minimizing the loss The gating network learns to adaptively weight the experts according to the problem characteristics, making its weighted output as close as possible to the discrimination probability of each expert. This collaborative learning mechanism enables the gating network and the expert model to adapt to each other and update iteratively during the training process, ultimately achieving an overall performance improvement of the model.

[0051] During training, a regularization term based on expert entropy is introduced to encourage the output of the gating network to have higher expert selectivity:

[0052]

[0053] Among them, H(·) represents the entropy function. Minimizing the regularization term R g Promotes the gating network to select as few experts as possible for each problem, avoids egalitarianism, and improves the utilization efficiency of experts.

[0054] Finally, the overall objective of the gating network is to minimize the unsupervised collaborative loss and the expert entropy regularization term. Among them, the gating loss function is:

[0055]

[0056] Among them, Represents the gating loss function, Represents the unsupervised loss, λ represents the predetermined balance weight, Represents the minimized regularization term. To sum up, the expert model realizes the discrimination of different types of values through supervised training of scientific value sentences. The gating network realizes an adaptive expert routing strategy through collaborative learning and entropy regularization. The expert model and the gating network cooperate with each other and optimize synergistically to form a multi-classification method for scientific value sentences.

[0057] Package the scientific value sentence recognition model and the scientific value sentence multi-classification model into a scientific value sentence detection dual model. Obtain a scientific and technological literature. Input the scientific and technological literature into the scientific value sentence detection dual model to obtain a comprehensive classification result of scientific value sentences.

[0058] The scientific value sentence recognition model and the scientific value sentence multi-classification model are encapsulated into a dual model for scientific value sentence detection. Subsequently, a scientific and technological literature is obtained. For example, a research paper on biodiversity is uploaded. The scientific and technological literature is input into the dual model for scientific value sentence detection, and through analysis by the dual model for scientific value sentence detection, a comprehensive classification result of scientific value sentences is obtained. This solves the technical problems of low processing efficiency and difficult accuracy guarantee in the existing method for classifying scientific literature value sentences when dealing with scientific and technological literature. By introducing a mixture of experts mechanism and data distribution consistency optimization, the classification accuracy of scientific value sentences can be effectively improved, achieving the technical effect of efficient and accurate analysis of multi-dimensional academic, application, and innovation values. Exemplarily, given an embedding representation x of a scientific question value sentence, three expert models f 1,2,3 (x) respectively output their discrimination scores in three value dimensions. The gating network g(x) calculates the weights of each expert according to the semantic features of the question, and finally obtains the comprehensive classification result:

[0059] y = g(x)1f1(x) + g(x)2f2(x) + g(x)3f3(x)

[0060] where f1(x), f2(x), and f3(x) respectively represent the outputs of the academic value, application value, and innovation value experts.

[0061] The method provided by the embodiment of this application further includes: locating the enrichment area of scientific value sentences based on the scientific and technological literature to obtain the enrichment area of value sentences. Inputting the enrichment area of value sentences into the scientific value sentence recognition model to obtain the recognition result of scientific value sentences. Inputting the recognition result of scientific value sentences into the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain multi-type value discrimination results. Inputting the multi-type value discrimination results into the gating network to generate the comprehensive classification result of scientific value sentences.

[0062] Input the scientific and technological literature into the dual model for detecting scientific value sentences to obtain the comprehensive classification result of scientific value sentences, including: In scientific and technological literature, knowledge is not evenly distributed, but shows a certain degree of concentration and regularity. Certain specific chapters or positions often contain a large amount of key information and core knowledge, with the characteristics of high knowledge density, large amount of information, and important content. These areas are called knowledge enrichment areas. Locate the knowledge enrichment areas of scientific value sentences based on the scientific and technological literature to obtain the value sentence enrichment areas. Input the value sentence enrichment areas into the scientific value sentence recognition model to obtain the scientific value sentence recognition result, and obtain sentences with scientific value. Input the scientific value sentence recognition result into the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain the multi-type value discrimination result. Finally, input the multi-type value discrimination result into the gating network to generate the comprehensive classification result of scientific value sentences. The comprehensive classification result of scientific value sentences is the comprehensive classification result of scientific value sentences belonging to academic value sentences, application value sentences, and innovation value sentences.

[0063] In the above text, with reference to Figure 1 A hybrid expert multi-classification method combined with an embedded dual model architecture according to an embodiment of the present invention is described in detail. Next, with reference to Figure 2 Describe a hybrid expert multi-classification system combined with an embedded dual model architecture according to an embodiment of the present invention.

[0064] The hybrid expert multi-classification system combined with an embedded dual model architecture according to an embodiment of the present invention solves the technical problems of low processing efficiency and difficult accuracy guarantee in the existing scientific literature value sentence classification method when dealing with scientific and technological literature. By introducing a hybrid expert mechanism and data distribution consistency optimization, it can effectively improve the classification accuracy of scientific value sentences and achieve the technical effect of efficient and accurate analysis of multi-dimensional academic, application, and innovation values. The hybrid expert multi-classification system combined with an embedded dual model architecture includes: a scientific value sentence recognition model construction module 11, a scientific value sentence multi-classification model construction module 12, a model encapsulation module 13, a scientific and technological literature acquisition module 14, and a comprehensive classification result acquisition module 15.

[0065] The scientific value sentence recognition model construction module 11 is used to perform deep learning on a pre-trained model based on a data distribution consistency threshold and a scientific value sentence training corpus to construct a scientific value sentence recognition model.

[0066] The scientific value sentence multi-classification model construction module 12 is used to build a scientific value sentence multi-classification model based on a hybrid expert mechanism, where the scientific value sentence multi-classification model includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network.

[0067] The model encapsulation module 13 is used to encapsulate the scientific value sentence recognition model and the scientific value sentence multi-classification model into a dual model for scientific value sentence detection.

[0068] The scientific literature acquisition module 14 is used to obtain scientific literature.

[0069] The comprehensive classification result acquisition module 15 is used to input the scientific literature into the dual model for scientific value sentence detection to obtain the comprehensive classification result of scientific value sentences.

[0070] Next, the specific configuration of the scientific value sentence recognition model construction module 11 will be described in detail. The scientific value sentence recognition model construction module 11 may further include: performing deep learning on a pre-trained model based on a data distribution consistency threshold and a scientific value sentence training corpus to construct a scientific value sentence recognition model, including: the scientific value sentence training corpus includes a scientific value sentence sample set and a non-scientific value sentence sample set. Divide the scientific value sentence training corpus according to a predetermined ratio to obtain a first corpus training set, a first corpus test set, and a first corpus validation set. Based on the data distribution consistency threshold, optimize the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set. Based on the second corpus training set, the second corpus test set, and the second corpus validation set, train, test, and validate the pre-trained model to generate the scientific value sentence recognition model.

[0071] Next, the specific configuration of the scientific value sentence recognition model construction module 11 will be further described in detail. The scientific value sentence recognition model construction module 11 further includes: based on the data distribution consistency threshold, optimizing the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set, including: calculating the positive and negative sample distributions of the first corpus training set, the first corpus test set, and the first corpus validation set respectively to obtain a first data distribution coefficient, a second data distribution coefficient, and a third data distribution coefficient. Based on the first data distribution coefficient, the second data distribution coefficient, and the third data distribution coefficient, evaluate the data distribution consistency of the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a data distribution consistency coefficient. Determine whether the data distribution consistency coefficient is greater than the data distribution consistency threshold. If the data distribution consistency coefficient is greater than the data distribution consistency threshold, adjust the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set according to the data distribution consistency threshold to generate the second corpus training set, the second corpus test set, and the second corpus validation set.

[0072] Next, the specific configuration of the multi-classification model construction module 12 for scientific value sentences will be described in detail. The multi-classification model construction module 12 for scientific value sentences may further include: building a multi-classification model for scientific value sentences based on a mixture of experts mechanism, including: obtaining multi-classification indicators for scientific value sentences, where the multi-classification indicators for scientific value sentences include academic value, application value, and innovation value. Based on the multi-classification indicators for scientific value sentences, loading a corpus for scoring academic value sentences, a corpus for scoring application value sentences, and a corpus for scoring innovation value sentences. Based on the corpus for scoring academic value sentences, training the academic value expert sub-model. Based on the corpus for scoring application value sentences, training the application value expert sub-model. Based on the innovation value expert sub-model, training the innovation value expert sub-model. Based on a gating loss function, constructing a gating network according to the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model. Using the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model as multi-classification parallel nodes. Connecting the multi-classification parallel nodes and the gating network to generate the multi-classification model for scientific value sentences.

[0073] Next, the specific configuration of the multi-classification model construction module 12 for scientific value sentences will be described in detail. The multi-classification model construction module 12 for scientific value sentences further includes: constructing a gating network according to the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model based on a gating loss function, including: collecting output samples of the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain an expert output sample set. Collecting discriminant probability parameters corresponding to the expert output sample set to obtain a discriminant probability distribution. Using minimizing the gating loss as the training objective of the gating network. Based on the gating loss function and the training objective of the gating network, performing unsupervised training on the expert output sample set and the discriminant probability distribution to generate the gating network.

[0074] Next, the specific configuration of the multi-classification model construction module 12 for scientific value sentences will be further described in detail. The multi-classification model construction module 12 for scientific value sentences further includes: The gating loss function is:

[0075]

[0076] where represents the gating loss function, represents the unsupervised loss, λ represents a predetermined balance weight, represents the minimization regularization term.

[0077] Next, the specific configuration of the scientific value sentence comprehensive classification result acquisition module 15 will be described in detail. The comprehensive classification result acquisition module 15 further includes: inputting the scientific and technological literature into the scientific value sentence detection dual model to obtain the scientific value sentence comprehensive classification result, including: locating the value sentence enrichment area based on the scientific and technological literature to obtain the value sentence enrichment area. Inputting the value sentence enrichment area into the scientific value sentence recognition model to obtain the scientific value sentence recognition result. Inputting the scientific value sentence recognition result into the academic value expert sub-model, the application value expert sub-model, and the innovation value expert sub-model to obtain the multi-type value discrimination result. Inputting the multi-type value discrimination result into the gating network to generate the scientific value sentence comprehensive classification result.

[0078] The hybrid expert multi-classification system provided by the embodiment of the present invention in combination with the embedded dual model architecture can execute the hybrid expert multi-classification method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0079] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The included individual units and modules are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized. In addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0080] The above specific implementation manners do not constitute a limitation to the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to the design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A hybrid expert multi-classification method combined with an embedded dual-model organizational structure, characterized in that, The method includes: Deeply learning the pre-trained model based on the data distribution consistency threshold and the training corpus of scientific value sentences to construct a scientific value sentence recognition model; Based on the mixture of experts mechanism, building a multi-classification model for scientific value sentences, where the multi-classification model for scientific value sentences includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model, and a gating network; Encapsulating the scientific value sentence recognition model and the multi-classification model for scientific value sentences into a dual model for scientific value sentence detection; Obtaining scientific and technological literature; Inputting the scientific and technological literature into the dual model for scientific value sentence detection to obtain the comprehensive classification result of scientific value sentences.

2. The method according to claim 1, wherein Deeply learning the pre-trained model based on the data distribution consistency threshold and the training corpus of scientific value sentences to construct a scientific value sentence recognition model, including: The training corpus of scientific value sentences includes a scientific value sentence sample set and a non-scientific value sentence sample set; Dividing the training corpus of scientific value sentences according to a predetermined ratio to obtain a first corpus training set, a first corpus test set, and a first corpus validation set; Based on the data distribution consistency threshold, optimizing the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set; Based on the second corpus training set, the second corpus test set, and the second corpus validation set, training, testing, and validating the pre-trained model to generate the scientific value sentence recognition model.

3. The method according to claim 2, wherein Based on the data distribution consistency threshold, optimizing the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a second corpus training set, a second corpus test set, and a second corpus validation set, including: Calculating the positive and negative sample distributions of the first corpus training set, the first corpus test set, and the first corpus validation set respectively to obtain a first data distribution coefficient, a second data distribution coefficient, and a third data distribution coefficient; Based on the first data distribution coefficient, the second data distribution coefficient, and the third data distribution coefficient, evaluating the data distribution consistency of the first corpus training set, the first corpus test set, and the first corpus validation set to obtain a data distribution consistency coefficient; Judging whether the data distribution consistency coefficient is greater than the data distribution consistency threshold; If the data distribution consistency coefficient is greater than the data distribution consistency threshold, adjusting the data distribution of the first corpus training set, the first corpus test set, and the first corpus validation set according to the data distribution consistency threshold to generate the second corpus training set, the second corpus test set, and the second corpus validation set.

4. The method according to claim 1, characterized in that, Based on the mixture of experts mechanism, building a multi-classification model for scientific value sentences, including: Obtaining multi-classification metrics for scientific value sentences, where the multi-classification metrics for scientific value sentences include academic value, application value, and innovation value; Based on the multi-classification metrics for scientific value sentences, loading the scoring corpus of academic value sentences, the scoring corpus of application value sentences, and the scoring corpus of innovation value sentences; Train the academic value expert sub-model based on the academic value sentence scoring corpus; Train the application value expert sub-model based on the application value sentence scoring corpus; Train the innovation value expert sub-model based on the innovation value expert sub-model; Construct the gating network based on the gating loss function according to the academic value expert sub-model, the application value expert sub-model and the innovation value expert sub-model; Use the academic value expert sub-model, the application value expert sub-model and the innovation value expert sub-model as multi-classification parallel nodes; Connect the multi-classification parallel nodes and the gating network to generate the scientific value sentence multi-classification model.

5. The method according to claim 4, characterized in that, Construct the gating network based on the gating loss function according to the academic value expert sub-model, the application value expert sub-model and the innovation value expert sub-model, including: Collect the output samples of the academic value expert sub-model, the application value expert sub-model and the innovation value expert sub-model to obtain an expert output sample set; Collect the discriminant probability parameters corresponding to the expert output sample set to obtain a discriminant probability distribution; Use minimizing the gating loss as the training objective of the gating network; Based on the gating loss function and the gating network training objective, perform unsupervised training on the expert output sample set and the discriminant probability distribution to generate the gating network.

6. The method according to claim 4, characterized in that, The gating loss function is: Among them, represents the gating loss function, represents the unsupervised loss, and λ represents a predetermined balance weight, represents minimizing the regularization term.

7. The method according to claim 1, wherein Input the scientific and technological literature into the scientific value sentence detection dual model to obtain the comprehensive classification result of the scientific value sentence, including: Locate the enrichment area of the scientific value sentence based on the scientific and technological literature to obtain the value sentence enrichment area; Input the value sentence enrichment area into the scientific value sentence recognition model to obtain the scientific value sentence recognition result; Input the scientific value sentence recognition result into the academic value expert sub-model, the application value expert sub-model and the innovation value expert sub-model to obtain multi-type value discrimination results; Input the multi-type value discrimination results into the gating network to generate the comprehensive classification result of the scientific value sentence.

8. A hybrid expert multi-classification system combined with an embedded dual-model organizational structure, characterized in that, The system includes: A scientific value sentence recognition model construction module, which is used to perform deep learning on a pre-trained model based on a data distribution consistency threshold and a scientific value sentence training corpus to construct a scientific value sentence recognition model; A scientific value sentence multi-classification model construction module, which is used to build a scientific value sentence multi-classification model based on the mixture of experts mechanism, wherein the scientific value sentence multi-classification model includes an academic value expert sub-model, an application value expert sub-model, an innovation value expert sub-model and a gating network; A model encapsulation module, which is used to encapsulate the scientific value sentence recognition model and the scientific value sentence multi-classification model into a scientific value sentence detection dual model; A scientific and technological literature acquisition module, which is used to obtain scientific and technological literature; A comprehensive classification result acquisition module, which is used to input the scientific and technological literature into the scientific value sentence detection dual model to obtain the comprehensive classification result of the scientific value sentence.

Citation Information

Patent Citations

  • Paper classification method based on graph neural network with mixed expert structure

    CN115510971A

  • Text processing method and device, equipment and medium

    CN116362240A