Classification method and device, electronic equipment and computer storage medium

By acquiring the data feature vectors of the source and target domains, and optimizing the model using the BERT network and cross loss function, the problems of pseudo-sample generation and catastrophic forgetting in discrete text data are solved, achieving effective transfer and accurate classification in unlabeled data.

CN120951077APending Publication Date: 2025-11-14DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510965901.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies are ineffective in generating pseudo-samples of discrete text data, and transfer methods based on deep pre-trained models are prone to catastrophic forgetting, affecting the model's performance in the target domain.

Method used

By acquiring labeled data from the source domain and unlabeled data from the target domain, target feature vectors are extracted, and the BERT network is used for training. The model is then optimized by combining a classification network and a domain discrimination network, and a cross-loss function is employed to mitigate catastrophic forgetting and uncover deep-seated common features between the source and target domains.

Benefits of technology

Even without labeled data, the model can effectively transfer and accurately classify data from different domains, mitigating catastrophic forgetting and improving the model's learning ability and classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951077A_ABST
    Figure CN120951077A_ABST
Patent Text Reader

Abstract

The invention relates to a classification method and device, electronic equipment and a computer storage medium, and the method comprises the steps: obtaining training data, the training data comprises source domain annotation data and target domain unannotation data, a source domain and a target domain are different fields, and the source domain annotation data comprises a plurality of training samples and a classification annotation result of each training sample; extracting a target feature vector of the training data; and training the initial network model according to the target feature vector to obtain a final classification model, and classifying the to-be-classified data of different domains based on the final classification model. According to the method provided by the invention, on the basis of the source domain annotated data of the existing annotated data, for a new to-be-migrated text, the disastrous forgetting is relieved through mixed training of the source domain annotated data and the target domain unannotated data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and more specifically, to a classification method, apparatus, electronic device, and computer storage medium. Background Technology

[0002] The current mainstream strategies for pseudo-sample generation and transfer in a certain field are: Strategy 1: Apply to the corresponding GAN network in this field, and expand the medical image training data by generating new sample images by setting generators and discriminators; Strategy 2: Model-based transfer learning, which involves fine-tuning and training a deep pre-trained model on target domain data to achieve transfer learning. Regarding Strategy 1, the following problems exist: Since GAN networks can only work on continuous real number domains and cannot be effective on discrete text data, there is currently no clear technical solution for generating pseudo samples for domains including discrete text data.

[0003] Regarding Strategy 2, the following problem exists: This common transfer method can cause catastrophic forgetting, that is, after the model is trained on the target domain data, its performance in the target domain improves, but the knowledge previously acquired in the source domain will be disturbed and suffer a significant decline.

[0004] Therefore, there is an urgent need in existing technologies for a solution that can alleviate the problem of catastrophic forgetting and improve the learning ability of models. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a classification method, apparatus, electronic device and computer storage medium, which aims to solve at least one of the above-mentioned technical problems.

[0006] In a first aspect, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a classification method, the method comprising: Acquire training data, which includes source domain labeled data and target domain unlabeled data. The source domain and target domain are different domains. The source domain labeled data includes multiple training samples and the classification labeling results of each training sample. Extract the target feature vector from the training data; The initial network model is trained based on the target feature vector to obtain the final classification model, which is then used to classify data in different domains.

[0007] The beneficial effects of this invention are as follows: Based on the source domain labeled data of the existing labeled data, when facing new text to be transferred (target domain unlabeled data), the mixed training of source domain labeled data and target domain unlabeled data alleviates catastrophic forgetting. Furthermore, training the initial network model based on the target feature vector can inherit the classification performance of the initial network model on the one hand, and deeply explore the deep common features between the source and target domains on the other hand, enabling the model to effectively transfer data without labeled data and alleviate the problem of catastrophic forgetting.

[0008] Based on the above technical solution, the present invention can be further improved as follows.

[0009] Furthermore, the aforementioned target feature vector includes a first feature vector corresponding to the labeled data in the source domain and a second feature vector corresponding to the unlabeled data in the target domain; the initial network model includes a classification network and a domain discrimination network. The initial network model is trained based on the target feature vector to obtain the final classification model, which includes: The classification network is trained based on the first feature vector to obtain the first classification prediction result; Based on the target feature vector, the domain discrimination network is trained to obtain the second classification prediction result of the input data source; The initial network model is trained based on the first and second classification prediction results to obtain the final classification model.

[0010] Furthermore, the initial network model is trained based on the first and second classification prediction results to obtain the final classification model, including: Based on the first classification prediction result and the classification labeling result corresponding to the source domain labeling data, the first loss value is determined by the first classification loss function; Based on the second classification prediction result and the true classification result of the domain to which each training sample belongs in the training data, the second loss value is determined by the second classification loss function. The initial network model is trained based on the first and second loss values ​​to obtain the final classification model.

[0011] Furthermore, the first classification loss function and the second classification loss function mentioned above are cross-loss functions.

[0012] Furthermore, the target feature vector extracted from the training data includes: The target feature vector is extracted from the training data using the BERT network.

[0013] Secondly, in order to solve the above-mentioned technical problems, the present invention also provides a sorting device, the device comprising: The acquisition module is used to acquire training data, which includes source domain labeled data and target domain unlabeled data. The source domain and target domain are different domains. The source domain labeled data includes multiple training samples and the classification labeling results of each training sample. The extraction module is used to extract the target feature vector from the training data; The training module is used to train the initial network model based on the target feature vector to obtain the final classification model, which is then used to classify the data to be classified in different domains.

[0014] Thirdly, in order to solve the above-mentioned technical problems, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the classification method of the present application.

[0015] Fourthly, in order to solve the above-mentioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the classification method of the present application.

[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below.

[0018] Figure 1 A flowchart illustrating a classification method according to an embodiment of the present invention; Figure 2 A schematic diagram of a basic framework provided for one embodiment of the present invention; Figure 3 This is a schematic diagram of a classification device provided in one embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation

[0019] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0020] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0021] The solution provided in this invention can be applied to any application scenario requiring cross-domain data migration training. The solution provided in this invention can be executed by any electronic device, such as a user's terminal device, including at least one of the following: smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart TV, or smart in-vehicle device.

[0022] This invention provides a possible implementation, such as... Figure 1 The diagram shows a flowchart of a classification method. This method can be executed by any electronic device, such as a terminal device, or jointly by a terminal device and a server. For ease of description, the method provided in this embodiment will be described below using a terminal device as the execution subject as an example. Figure 1 The flowchart shown indicates that the method may include the following steps: S10, Obtain training data. The training data includes source domain labeled data and target domain unlabeled data. The source domain and target domain are different domains. The source domain labeled data includes multiple training samples and the classification labeling results of each training sample. S20, extract the target feature vector from the training data; S30: Train the initial network model based on the target feature vector to obtain the final classification model, and classify the data to be classified in different domains based on the final classification model.

[0023] By employing the method of this invention, based on existing labeled source domain data, and facing new text to be transferred (unlabeled target domain data), the method mitigates catastrophic forgetting by training the model with a mixture of labeled source domain data and unlabeled target domain data. Furthermore, by training the initial network model based on the target feature vector, the model can inherit the classification performance of the initial network model while also deeply exploring the common features between the source and target domains. This allows the model to effectively transfer data even without labeled data, thus mitigating the problem of catastrophic forgetting.

[0024] The present invention will be further described below with reference to the following specific embodiments. In this embodiment, a classification method may include the following steps: S10, Obtain training data. The training data includes source domain labeled data and target domain unlabeled data. The source domain and target domain are different domains. The source domain labeled data includes multiple training samples and the classification labeling results of each training sample. Here, source domain labeled data refers to the training data corresponding to the source domain, and target domain unlabeled data refers to the training data corresponding to the target domain. The training samples in the target domain unlabeled data are unlabeled training samples.

[0025] Optionally, the source domain can be a medical classification domain, a vehicle classification domain, etc. If the source domain is a medical classification domain, the classification result of the corresponding training sample can be a disease diagnosis result or a classification result of the department to which the disease belongs. If the source domain is a vehicle classification domain, the classification result of the corresponding training sample can be a vehicle type classification result.

[0026] S20, extract the target feature vector from the training data; Optionally, the target feature vector of the training data can be extracted using the BERT network. Other feature extraction networks or models, such as RoBERTa, ALBERT, GPT, etc., can also be considered, or combined with domain-specific feature extractors, such as BioBERT in the medical field.

[0027] Specifically, the aforementioned target feature vector includes a first feature vector corresponding to the source domain labeled data and a second feature vector corresponding to the target domain unlabeled data.

[0028] The initial network model described above includes a classification network and a domain discriminant network. Training the initial network model based on the target feature vector yields the final classification model, which includes: The classification network is trained based on the first feature vector to obtain the first classification prediction result. The first feature vector includes the feature vector corresponding to each training sample in the source domain labeled data. Each training sample in the source domain labeled data corresponds to a classification result. Therefore, the first classification prediction result includes the classification results corresponding to all training samples in the source domain labeled data.

[0029] Based on the target feature vector, the domain discrimination network is trained to obtain the second classification prediction result of the input data source; wherein, the second feature vector includes the feature vector corresponding to each training sample in the unlabeled data of the target domain, and each training sample in the unlabeled data of the target domain corresponds to a classification result, so the second classification prediction result includes the classification results corresponding to all training samples in the unlabeled data of the target domain.

[0030] The initial network model is trained based on the first and second classification prediction results to obtain the final classification model.

[0031] The classification network can be an FCN fully connected network, and the domain discrimination network can also be an FCN fully connected network.

[0032] Optionally, the initial network model is trained based on the first and second classification prediction results to obtain the final classification model, including: Based on the first classification prediction result and the classification labeling result corresponding to the source domain labeling data, the first loss value is determined by the first classification loss function; Based on the second classification prediction result and the true classification result (source domain or target domain) of each training sample in the training data, the second loss value is determined by the second classification loss function. The initial network model is trained based on the first and second loss values ​​to obtain the final classification model.

[0033] The first classification prediction result can be characterized by a category identifier, while the second classification prediction result can be represented by different identifiers, such as 0 and 1, where 0 represents the source domain and 1 represents the target domain.

[0034] By working together with two networks (a classification network and a domain discriminant network), the classification network d can make predictions normally, while the domain discriminant network j cannot identify whether the vector extracted by the BERT network comes from the source domain or the target domain. This indicates that the model has captured the deep common features between different domains.

[0035] Optionally, the first classification loss function and the second classification loss function mentioned above are cross-loss functions.

[0036] Optionally, one implementation of determining the first loss value using the first classification loss function based on the first classification prediction result and the classification labeling result corresponding to the source domain labeled data is as follows: Where loss_d is the first loss value, P d (v) and q d (v) represents the probability distribution of the predicted value (first classification prediction result) and the actual value (classification labeling result corresponding to the source domain labeled data) obtained by the domain discriminant network d for the input vector v (first feature vector).

[0037] Optionally, one implementation of determining the second loss value using the second classification loss function based on the second classification prediction result and the true classification result (source domain or target domain) of each training sample in the training data is as follows: Where loss_j is the second loss value, P j (v) and qj (v) represents the probability distribution of the predicted value (second classification prediction result) and the actual value (true classification result of the domain to which each training sample belongs in the training data) obtained by the classification network j for the input vector v (target feature vector).

[0038] Alternatively, one implementation of training the initial network model based on the first loss value and the second loss value to obtain the final classification model is as follows: The total loss value is determined based on the first loss value and the second loss value. The initial network model is trained based on the total loss value to obtain the final classification model.

[0039] Optionally, one way to determine the total loss value based on the first loss value and the second loss value is as follows: Loss_all = loss_j - loss_d = Here, Loss_all represents the total loss value.

[0040] By continuously training and reducing the total loss value Loss_all, the model eventually converges to obtain the final classification model.

[0041] S30: Train the initial network model based on the target feature vector to obtain the final classification model, and classify the data to be classified in different domains based on the final classification model.

[0042] Based on the solution proposed in this application, even when no labeled data is required for a new target domain, the features in the unlabeled data of the new target domain can still be effectively transferred based on the solution proposed in this application. This allows the model to learn the features in the data of different domains more accurately, so that the final classification model can classify the data of different domains more accurately.

[0043] In this application, a dynamic feature extraction mechanism can be introduced to adaptively select the optimal feature extraction network based on the different characteristics of source and target domain data.

[0044] The core of the dynamic feature extraction mechanism lies in adaptively selecting the optimal feature extractor based on the characteristics of the source and target domain data. This mechanism can dynamically adjust the structure and parameters of the feature extractor according to the data type (such as text, image, audio, etc.), the domain characteristics of the data (such as medicine, finance, transportation, etc.), and the quality of the data (such as noise level, data integrity, etc.), thereby achieving efficient feature extraction for different types of data.

[0045] In this application, a feature extractor selector can be preset to dynamically select the optimal feature extractor based on the characteristics of the input data. The specific selection logic is as follows: Data type identification: The data type (text, image, audio, etc.) is initially determined by the data's metadata (such as file format, data dimension, etc.).

[0046] Domain-specific characteristics analysis: Utilizing domain-specific labels or pre-trained models to analyze the domain characteristics of data (such as medicine, finance, etc.).

[0047] Data quality assessment: Evaluating the quality of data, including noise level, data integrity, etc.

[0048] Performance evaluation: Based on historical data and experimental results, evaluate the performance of different feature extractors on similar data, and select the feature extractor with the best performance.

[0049] Adaptive adjustment: The feature extractor selector can dynamically adjust the selection logic based on feedback during training to adapt to new data characteristics.

[0050] In some cases, a single feature extractor may not be able to fully capture the features of the data. Therefore, a feature fusion module can be built to combine the outputs of multiple feature extractors.

[0051] The specific fusion strategy is as follows: Weighted average: The outputs of multiple feature extractors are weighted according to their performance weights.

[0052] Cascaded fusion: The outputs of multiple feature extractors are sequentially input into the next feature extractor to form a cascaded structure.

[0053] Multimodal fusion: For multimodal data (such as data containing both text and images), design specialized multimodal fusion strategies, such as attention-guided fusion.

[0054] Dynamic adjustment: The feature fusion module can dynamically adjust the fusion strategy based on performance feedback during training to achieve the optimal fusion effect.

[0055] Based on the above principles, a further solution is as follows: 1. Input data preprocessing: The input data undergoes preliminary processing to extract metadata (such as file format and data dimensions) and performs data quality assessment. Based on data type and domain characteristics, the data is divided into different subsets, each corresponding to a different feature extractor.

[0056] 2. Feature Extractor Selection: The feature extractor selector chooses the optimal feature extractor from the pool of feature extractors based on the characteristics of the input data.

[0057] If a single feature extractor cannot meet the requirements, the feature extractor selector can select multiple feature extractors and pass their outputs to the feature fusion module.

[0058] 3. Feature extraction and fusion: The selected feature extractor extracts features from the input data and generates the target feature vector.

[0059] If multiple feature extractors are used, the feature fusion module merges these feature vectors to generate the final target feature vector.

[0060] 4. Model Training and Optimization: The fused feature vectors are then input into the initial feature extraction network for training.

[0061] During training, the logic of the feature extractor selector and the strategy of the feature fusion module are dynamically adjusted based on the model's performance feedback to optimize the model's performance.

[0062] As an example, in the medical field, training data can include different types of data. Common data types include electronic medical records (text data), X-ray images (image data), and electrocardiograms (audio data). These data come from different source and target domains, and their quality and characteristics vary. For example, electronic medical records may have inconsistent text formats, X-ray images may have varying noise levels, and electrocardiograms may have inconsistent signal integrity.

[0063] The training process of the feature extraction network is as follows: 1. Construct a feature extractor pool based on different types of data, specifically: Text data feature extractors: BioBERT, RoBERTa.

[0064] Image data feature extractors: ResNet50, InceptionV3.

[0065] Audio data feature extractors: Wav2Vec2.0, DeepSpeech.

[0066] 2. Set the feature extractor selector: Data type identification: Determine the data type based on file format and data dimensions.

[0067] Domain-specific characteristics analysis: Utilize pre-trained models in the medical field (such as BioBERT) to analyze the medical domain characteristics of the data.

[0068] Data quality assessment: Evaluate the format consistency of text data, the noise level of image data, and the signal integrity of audio data.

[0069] Performance evaluation: Based on historical data and experimental results, evaluate the performance of different feature extractors on similar data, and select the feature extractor with the best performance.

[0070] 3. Configure the feature fusion module: For multimodal data (such as data that simultaneously includes electronic medical records, X-rays, and electrocardiograms), an attention-guided multimodal fusion strategy is used to fuse feature vectors from different modalities.

[0071] 4. Model Training and Optimization: The fused feature vectors are input into the initial feature extraction network for training. During training, the logic of the feature extractor selector and the strategy of the feature fusion module are dynamically adjusted based on the model's performance feedback to optimize model performance.

[0072] Through a dynamic feature extraction mechanism, this model demonstrates outstanding performance in multi-domain transfer learning within the medical field. When the source domain is electronic medical record data from one hospital and the target domain is electronic medical record data from another hospital, the model can dynamically select the optimal feature extractor based on the characteristics and quality of the data, significantly improving the effectiveness of transfer learning. Furthermore, when processing multimodal data, through a multimodal fusion strategy, the model can effectively extract and fuse features from different modalities, improving the accuracy of disease diagnosis.

[0073] To better illustrate and understand the principle of the method provided by this invention, the following description uses an optional specific embodiment to illustrate the solution of this invention. It should be noted that the specific implementation of each step in this specific embodiment should not be construed as a limitation of the solution of this invention. Other implementations that can be conceived by those skilled in the art based on the principle of the solution provided by this invention should also be considered within the scope of protection of this invention.

[0074] See Figure 2 The basic framework shown, taking the medical diagnostics field as an example, will be further described below: Data in medical settings has the following characteristics: Privacy is paramount; patient data must be anonymized before it can be used for model training.

[0075] It is highly specialized, the data is very specialized, and annotation is difficult, requiring a high level of skill and expertise from the annotators.

[0076] Due to their proprietary nature, different hospitals have different standards for writing electronic medical records, and different doctors have different writing styles.

[0077] Therefore, obtaining a large amount of high-quality medical labeled data requires a significant investment of human and material resources. Moreover, even if a diagnostic model has been trained for one or a few departments, when it is applied to real-world scenarios, especially when the data is cross-departmental or cross-hospital, the new data needs to be re-labeled, resulting in huge time and financial costs.

[0078] To address this issue, this application proposes a method that combines source domain labeled data and target domain unlabeled data with adversarial networks to extract and retain deep common features of data from different departments / hospitals, thereby reducing manual annotation and solving the problem of cross-domain diagnostic models.

[0079] In this embodiment, for source domain S, the source domain labeled data is electronic medical record text, the content of which is: "Currently 2 years and 8 months old, with poor language expression, hyperactivity, poor concentration, childish expression, able to socialize, lively, normal activity, average motor coordination, and normal eating." The corresponding classification labeling result, i.e., the diagnosis result, is: Language development disorder.

[0080] The data processing procedure for this application is as follows: Step 1. Mix the training samples from the source domain S and the target domain T, and label them (S,T) according to their respective domains to obtain the training data; Step 2. Feed the training data into BERT for feature extraction to obtain vectors V_s and V_t, where V_s is the first feature vector and V_t is the second feature vector; Step 3. All feature vectors (V_s, V_t) are trained through the domain discriminant module J (i.e., the domain discriminant network J) to obtain the prediction result (second classification prediction result) for the input data source. The second loss value is determined by loss_j. The output of the domain discriminant module J can be 0 or 1, mapping the data source category S / T, and the data source category (i.e., the second classification prediction result of the domain discriminant module, the result is S or T).

[0081] The vector labeled V_s, originating from the source domain S, is trained using the diagnostic classification module D (i.e., the classification network) to obtain a diagnostic prediction result (the first classification prediction result). The first loss value is determined by loss_d. The output of the diagnostic classification module D can be 0, 1, 2..., representing the number of diseases in the labeled data and corresponding to the number of output values, which are mapped to disease diagnosis categories. Specifically, the first classification prediction result (i.e., the first classification prediction result of the diagnostic classification module) can represent the specific disease.

[0082] Step 4. Through comprehensive training, reduce the loss_d of the diagnostic classification module and increase the loss_j of the domain discrimination module to train and obtain the final classification model.

[0083] The solution of the present invention has the following beneficial effects: Building upon existing labeled data (source domain), when faced with new text to be transferred (target domain text), catastrophic forgetting is mitigated by training with a mix of source and target domain data. Furthermore, the combined effect of two modules (diagnostic classification module and domain discriminant module) ensures that the diagnostic classification module D can perform diagnostic predictions correctly, while the domain discriminant module J cannot distinguish whether the vectors extracted by the BERT feature extraction module come from the source or target domain (indicating that the model has captured the deep common features between different domains). This achieves the dual goals of inheriting the original model's diagnostic classification performance while deeply exploring the common features between the source and target domains, enabling the model to effectively transfer knowledge even without labeled data and mitigating the problem of catastrophic forgetting.

[0084] Based on and Figure 1 Based on the same principle as the method shown, this embodiment of the invention also provides a classification device 20, such as... Figure 3 As shown, the classification device 20 may include an acquisition module 210, an extraction module 220, and a training module 230, wherein: The acquisition module 210 is used to acquire training data, which includes source domain labeled data and target domain unlabeled data. The source domain and target domain are different domains. The source domain labeled data includes multiple training samples and the classification labeling results of each training sample. Extraction module 220 is used to extract the target feature vector from the training data; The training module 230 is used to train the initial network model based on the target feature vector to obtain the final classification model, so as to classify the data to be classified in different domains based on the final classification model.

[0085] Optionally, the aforementioned target feature vector includes a first feature vector corresponding to the source domain labeled data and a second feature vector corresponding to the target domain unlabeled data; the initial network model includes a classification network and a domain discrimination network, and the training module 230, when training the initial network model based on the target feature vector to obtain the final classification model, is specifically used for: The classification network is trained based on the first feature vector to obtain the first classification prediction result; Based on the target feature vector, the domain discrimination network is trained to obtain the second classification prediction result of the input data source; The initial network model is trained based on the first and second classification prediction results to obtain the final classification model.

[0086] Optionally, when the training module 230 trains the initial network model based on the first classification prediction result and the second classification prediction result to obtain the final classification model, it is specifically used for: Based on the first classification prediction result and the classification labeling result corresponding to the source domain labeling data, the first loss value is determined by the first classification loss function; Based on the second classification prediction result and the true classification result of the domain to which each training sample belongs in the training data, the second loss value is determined by the second classification loss function. The initial network model is trained based on the first and second loss values ​​to obtain the final classification model.

[0087] Optionally, the first classification loss function and the second classification loss function mentioned above are cross-loss functions.

[0088] Optionally, when extracting the target feature vector from the training data, the extraction module 220 is specifically used for: The target feature vector is extracted from the training data using the BERT network.

[0089] The classification device of this invention can execute the classification method provided in this invention, and their implementation principles are similar. The actions performed by each module and unit in the classification device in each embodiment of this invention correspond to the steps in the classification method in each embodiment of this invention. For detailed functional descriptions of each module of the classification device, please refer to the descriptions in the corresponding classification methods shown above, which will not be repeated here.

[0090] The aforementioned classification device may be a computer program (including program code) running on a computer device, such as an application software; the device may be used to execute the corresponding steps in the method provided in the embodiments of the present invention.

[0091] In some embodiments, the classification device provided in this invention can be implemented using a combination of hardware and software. As an example, the classification device provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the classification method provided in this invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0092] In other embodiments, the classification device provided in this invention can be implemented in software. Figure 3A classification device stored in a memory is shown, which may be software in the form of programs and plug-ins, and includes a series of modules, including an acquisition module 210, an extraction module 220 and a training module 230, for implementing the classification method provided in the embodiments of the present invention.

[0093] The modules described in the embodiments of the present invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0094] Based on the same principles as the methods shown in the embodiments of the present invention, the embodiments of the present invention also provide an electronic device, which may include, but is not limited to: a processor and a memory; the memory for storing computer programs; and the processor for executing the methods shown in any embodiment of the present invention by invoking the computer programs.

[0095] In one alternative embodiment, an electronic device is provided, such as Figure 4 As shown, Figure 4 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0096] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0097] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0098] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0099] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0100] Among these, electronic devices can also be terminal devices. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0101] This invention provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0102] According to another aspect of the present invention, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.

[0103] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0104] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0105] The computer-readable storage medium provided in this invention can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0106] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0107] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

Claims

1. A classification method, characterized in that, include: Acquire training data, which includes source domain labeled data and target domain unlabeled data. The source domain and target domain are different domains. The source domain labeled data includes multiple training samples and the classification labeling results of each training sample. Extract the target feature vector from the training data; The initial network model is trained based on the target feature vector to obtain the final classification model, which is then used to classify data to be classified in different domains.

2. The method according to claim 1, characterized in that, The target feature vector includes a first feature vector corresponding to the source domain labeled data and a second feature vector corresponding to the target domain unlabeled data; the initial network model includes a classification network and a domain discrimination network, and the step of training the initial network model based on the target feature vector to obtain the final classification model includes: The classification network is trained based on the first feature vector to obtain a first classification prediction result; Based on the target feature vector, the domain discrimination network is trained to obtain the second classification prediction result of the input data source; The initial network model is trained based on the first classification prediction result and the second classification prediction result to obtain the final classification model.

3. The method according to claim 2, characterized in that, The step of training the initial network model based on the first classification prediction result and the second classification prediction result to obtain the final classification model includes: Based on the first classification prediction result and the classification labeling result corresponding to the source domain labeling data, a first loss value is determined by a first classification loss function; Based on the second classification prediction result and the true classification result of the domain to which each training sample belongs in the training data, the second loss value is determined by the second classification loss function. The initial network model is trained based on the first loss value and the second loss value to obtain the final classification model.

4. The method according to claim 3, characterized in that, The first classification loss function and the second classification loss function are cross loss functions.

5. The method according to any one of claims 1 to 4, characterized in that, The extraction of the target feature vector from the training data includes: The target feature vector of the training data is extracted using the BERT network.

6. A sorting device, characterized in that, include: The acquisition module is used to acquire training data, which includes source domain labeled data and target domain unlabeled data. The source domain and target domain are different domains. The source domain labeled data includes multiple training samples and the classification labeling results of each training sample. The extraction module is used to extract the target feature vector from the training data; The training module is used to train the initial network model based on the target feature vector to obtain the final classification model, so as to classify the data to be classified in different domains based on the final classification model.

7. The apparatus according to claim 6, characterized in that, The target feature vector includes a first feature vector corresponding to the source domain labeled data and a second feature vector corresponding to the target domain unlabeled data; the initial network model includes a classification network and a domain discrimination network; when the training module trains the initial network model based on the target feature vector to obtain the final classification model, it is specifically used for: The classification network is trained based on the first feature vector to obtain a first classification prediction result; Based on the target feature vector, the domain discrimination network is trained to obtain the second classification prediction result of the input data source; The initial network model is trained based on the first classification prediction result and the second classification prediction result to obtain the final classification model.

8. The apparatus according to claim 7, characterized in that, When the training module trains the initial network model based on the first classification prediction result and the second classification prediction result to obtain the final classification model, it is specifically used for: Based on the first classification prediction result and the classification labeling result corresponding to the source domain labeling data, a first loss value is determined by a first classification loss function; Based on the second classification prediction result and the true classification result of the domain to which each training sample belongs in the training data, the second loss value is determined by the second classification loss function. The initial network model is trained based on the first loss value and the second loss value to obtain the final classification model.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method of any one of claims 1-5.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-5.