Cross-domain semantic matching method and system for mechanical and electrical products
Through incremental pre-training and data augmentation methods, a cross-domain semantic matching model is constructed, which solves the problems of low generalization and accuracy of cross-domain knowledge retrieval models for electromechanical products, and achieves more efficient cross-domain semantic matching and knowledge retrieval.
Patent Information
- Application Number
- CN202510408514.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing cross-domain knowledge retrieval model of mechanical and electrical products has low accuracy and poor generalization in natural language query, making it difficult to effectively capture the deep semantic associations of cross-domain texts.
By obtaining samples related to electromechanical products in multiple fields, using pre-trained models for incremental pre-training and large language model data enhancement, combining labeled samples and enhanced samples for supervision and fine-tuning, a cross-domain semantic matching model is constructed.
It improves the generalization and accuracy of the cross-domain semantic matching model, can better understand cross-domain natural language query and match relevant knowledge, and improves the retrieval accuracy.
Smart Images

Figure CN120256603A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a method and system for cross-domain semantic matching of electromechanical products. Background Art
[0002] The semantic matching model is an artificial intelligence model based on deep learning, which aims to determine the semantic relevance between texts by analyzing the semantic features of text data. Existing innovative designs of electromechanical products rely on cross-domain knowledge (such as patents, scientific effects, and material science), but traditional retrieval models (such as methods based on keyword matching or static word embedding) are difficult to capture the deep semantic association between natural language queries and cross-domain texts. Therefore, the cross-domain knowledge retrieval in existing electromechanical product design has low accuracy for natural language (unstructured) queries and poor generalization of cross-domain retrieval models. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a method and system for cross-domain semantic matching of electromechanical products, so as to improve the problem of poor generalization of cross-domain semantic matching models of electromechanical products.
[0004] In a first aspect, the present application provides a method for cross-domain semantic matching of electromechanical products, comprising: obtaining a sample set, the sample set comprising sub-samples related to electromechanical products in multiple fields, the sub-samples comprising labeled samples and unlabeled samples, wherein the labeled samples comprise query statements, text data, and a correlation index between the query statements and the text data; performing incremental pre-training on a pre-trained model according to the unlabeled samples to obtain an intermediate model; performing data enhancement on the labeled samples using a large language model to obtain enhanced samples; and performing supervised fine-tuning on the intermediate model according to the labeled samples and the enhanced samples to obtain a cross-domain semantic matching model.
[0005] In the above scheme, using sub-samples of multiple fields related to electromechanical products for training can tap into the potential connections between knowledge in different fields, promote the deep integration of knowledge, and enable the cross-domain semantic matching model to be exposed to the language expressions and knowledge structures of different fields, thereby enhancing its adaptability to diversified data and optimizing the generalization of the model. Unlabeled samples are used in the incremental pre-training stage to allow the cross-domain semantic matching model to learn a wider range of semantic features and patterns, so that it can better understand the query intent and match relevant knowledge when faced with cross-domain natural language queries. In the supervised fine-tuning stage, the generalization of the cross-domain semantic matching model is further optimized by combining labeled samples and enhanced samples, so that the model can more accurately judge the relevance of query statements and text data when retrieving cross-domain knowledge of electromechanical products, thereby improving the retrieval accuracy.
[0006] As an alternative approach, the incremental pre-training of the pre-trained model based on the unlabeled samples includes: based on the embedding vectors of the unlabeled samples generated by the pre-trained model, constructing the first positive sample pairs and the first negative sample pairs using the SimCSE contrastive learning method; inputting the similarities of the first positive sample pairs and the similarities of the first negative sample pairs into the InfoNCE loss function to calculate the loss value; and optimizing the parameters of the pre-trained model by minimizing the loss value of the InfoNCE loss function.
[0007] In the above solution, by using SimCSE contrastive learning to construct the first positive and negative sample pairs based on the embedding vectors of the unlabeled samples generated by the pre-trained model, and then calculating the loss value through the InfoNCE loss function and optimizing the model parameters, it is possible to fully exploit the value of the unlabeled data, improve the cross-domain semantic matching model's ability to judge semantic similarity, enhance the model's generalization performance, optimize the training effect, enable the model to better adapt to the cross-domain semantic matching task of electromechanical products, and efficiently and accurately judge the text semantic association.
[0008] As an alternative approach, the constructing of the first positive sample pairs and the first negative sample pairs using the SimCSE contrastive learning method based on the embedding vectors of the unlabeled samples generated by the pre-trained model includes: randomly sampling a preset number of historical embedding vectors from the experience replay buffer in each batch of incremental pre-training; constructing the first positive sample pairs and the first negative sample pairs based on the historical embedding vectors and the current embedding vectors; and the optimizing of the parameters of the pre-trained model by minimizing the loss value of the InfoNCE loss function includes: calculating the gradient of the pre-trained model parameters according to the InfoNCE loss function; and pruning the gradient to keep the gradient within a preset range, where the gradient is used to optimize the parameters of the pre-trained model.
[0009] In the above solution, sampling historical embedding vectors from the experience replay buffer during incremental pre-training and constructing positive and negative sample pairs with the current embedding vectors increases the sample diversity, helps the cross-domain semantic matching model consolidate old knowledge and reduce catastrophic forgetting. At the same time, pruning the gradient calculated by the InfoNCE loss function to keep the gradient within a preset range avoids gradient explosion or disappearance, makes the model training more stable, and also reduces catastrophic forgetting, ultimately improving the performance of the pre-trained model in the cross-domain semantic matching task.
[0010] As an alternative approach, the data augmentation of the labeled samples using the large language model includes: sending at least one instruction to the large language model to perform data augmentation on the labeled samples, where the instruction is used to instruct the large language model to output text data with the same semantics as the labeled samples but different expressions.
[0011] In the above solution, by sending instructions to the large language model to perform data augmentation on the labeled samples, text data with the same semantics as the labeled samples but different expression forms is obtained, increasing the diversity of the training data. This enables the cross-domain semantic matching model to learn more diverse language expression patterns during supervised fine-tuning, helps improve the model's semantic understanding ability under different expressions, and further enhances the generalization and accuracy of the model in cross-domain semantic matching tasks.
[0012] As an alternative, the supervised fine-tuning of the intermediate model according to the labeled samples and the augmented samples includes: calculating the cosine similarity between the query statement and the corresponding text data in the labeled samples and the cosine similarity between the query statement and the corresponding text data in the augmented samples; scaling the cosine similarity between the query statement and the text data corresponding to the relevance index according to the relevance index to obtain a scaled cosine similarity; and performing supervised fine-tuning on the intermediate model based on the scaled cosine similarity.
[0013] In the above solution, by calculating the cosine similarity between the query statement and the text data in the labeled samples and the augmented samples, scaling the similarity according to the relevance index to obtain a scaled cosine similarity to adjust the weights of difficult samples during training, and then performing supervised fine-tuning on the intermediate model based on this, the fine-tuned cross-domain semantic matching model can more accurately capture semantic matching relationships, strengthen the learning of difficult samples, improve the model's semantic matching ability on different data, optimize the model performance, and improve the accuracy and efficiency of retrieval.
[0014] As an alternative, the labeled samples and the augmented samples include second positive sample pairs and second negative sample pairs, where the second positive sample pairs are query statements and text data determined to be highly relevant according to the relevance index and a preset rule, and the second negative sample pairs are query statements and text data determined to be lowly relevant according to the relevance index and the preset rule.
[0015] In the above solution, by clarifying the second positive sample pairs and the second negative sample pairs in the labeled samples and the augmented samples, that is, dividing the query statements and text data with high and low relevance according to the relevance index and a preset rule, it provides a clear positive and negative sample reference for the model's supervised fine-tuning, enabling the cross-domain semantic matching model to more targeted learn the key features of semantic matching during the training process, strengthen the ability to distinguish between similar and dissimilar text pairs, and effectively improve the accuracy and reliability of the model in cross-domain semantic matching tasks.
[0016] As an alternative, the supervised fine-tuning of the intermediate model based on the scaled cosine similarity includes: determining a loss value according to the scaled cosine similarity and a loss function, and adjusting the parameters of the intermediate model according to the loss value to obtain the cross-domain semantic matching model.
[0017] In the above solution, determining the loss value based on the scaled cosine similarity and the loss function, and accordingly adjusting the parameters of the intermediate model to obtain the cross-domain semantic matching model can use the scaled cosine similarity to more accurately reflect the semantic matching degree between the query statement and the text data, make the loss value more accurately measure the difference between the model prediction and the actual situation, guide the model to optimize the parameters, thereby improving the performance of the model in cross-domain semantic matching, enhancing the accuracy and efficiency of semantic matching, and promoting the effective retrieval and application of cross-domain knowledge.
[0018] As an alternative, the loss function is the CoSENT function with dynamic weights added, and the expression of the CoSENT function is as follows: wherein, the is the loss value of the CoSENT function, the i, the j, and the sim(i, j) are respectively the query statement, the text data, and the correlation index of the second positive sample pair, the k, the l, and the sim(k, l) are respectively the query statement, the text data, and the correlation index of the second negative sample pair, the λ is a hyperparameter, the ω kl is the weight factor of the second negative sample pair, the ω ij is the weight factor of the second positive sample pair, the cos(u k , u l ) is the predicted similarity of the second negative sample pair, and the cos(u i , u j ) is the predicted similarity of the second positive sample pair.
[0019] In the above solution, using the CoSENT function with dynamic weights added as the loss function, adjusting the scale of the loss function through the hyperparameter λ, and dynamically adjusting the importance of the sample pairs in the loss calculation according to the weight factors of the second positive and negative sample pairs, enabling the model to flexibly adjust the learning focus according to the predicted similarity of the sample pairs, assigning higher weights to difficult sample pairs, thereby enhancing the model's ability to distinguish the second positive and negative sample pairs, effectively improving the accuracy of the model in calculating semantic similarity, optimizing the performance in the cross-domain semantic matching task, and better coping with complex and diverse semantic matching scenarios.
[0020] In a second aspect, the present application provides an electronic device, including: a processor, a memory, and a bus. Among them, the processor and the memory complete communication with each other through the bus; the memory stores program instructions executable by the processor, and the processor can execute the method steps of the first aspect by invoking the program instructions.
[0021] In a third aspect, the present application provides a cross-domain semantic matching device for electromechanical products, including: an acquisition module for acquiring a sample set, the sample set including multiple sub-samples related to the electromechanical products within multiple fields, the sub-samples including labeled samples and unlabeled samples. Among them, the labeled samples include query statements, text data, and a correlation index between the query statements and the text data; an incremental pre-training module for performing incremental pre-training on a pre-trained model according to the unlabeled samples to obtain an intermediate model; a data augmentation module for using a large language model to perform data augmentation on the labeled samples to obtain augmented samples; a supervised fine-tuning module for performing supervised fine-tuning on the intermediate model according to the labeled samples and the augmented samples to obtain a cross-domain semantic matching model.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium, including: the computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method steps of the first aspect.
[0023] In a fifth aspect, the present application provides a computer program product, including computer program instructions, which, when read and run by a processor, execute the method steps of the first aspect.
[0024] In a sixth aspect, the present application provides a system on which a cross-domain semantic matching model is deployed; the cross-domain semantic matching model is constructed by using the method steps of the first aspect.
[0025] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will become apparent from the specification, or will be understood by implementing the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is a schematic flowchart of a cross-domain semantic matching method for electromechanical products provided by an embodiment of the present application;
[0028] Figure 2 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application;
[0029] Figure 3 Schematic diagram of the structure of a cross - domain semantic matching device for electromechanical products provided by the embodiment of the present application. Specific embodiments
[0030] Hereinafter, embodiments of the technical solution of the present application will be described in detail with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solution of the present application more clearly, and thus are only examples and should not be used to limit the protection scope of the present application.
[0031] It should be noted that all the technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above - mentioned drawings are intended to cover non - exclusive inclusion.
[0032] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary - secondary relationship of the indicated technical features. In the description of the embodiments of the present application, "a plurality of" means more than two unless otherwise specifically defined.
[0033] It can be understood that the cross - domain semantic matching method for electromechanical products provided by the embodiments of the present application can be applied to terminal devices (also referred to as electronic devices) and servers; among them, the terminal device can specifically be a smart phone, a tablet computer, a computer, a personal digital assistant (PDA), etc.; the server can specifically be an application server or a Web server.
[0034] For the convenience of understanding the technical solution provided by the embodiments of the present application, hereinafter, taking the server as the execution subject as an example, the application scenario of the cross - domain semantic matching method for electromechanical products provided by the embodiments of the present application will be introduced.
[0035] Refer to Figure 1 , Figure 1 Schematic flowchart of a cross - domain semantic matching method for electromechanical products provided by the embodiments of the present application. The method includes the following steps:
[0036] Step S10: Obtain a sample set, which includes sub-samples related to electromechanical products in multiple fields. The sub-samples include labeled samples and unlabeled samples. Among them, the labeled samples include query statements, text data, and the correlation index between the query statement and the text data.
[0037] To construct a sample set covering multi-field knowledge, the sample set includes sub-samples related to electromechanical products in multiple fields. When collecting training data related to electromechanical products in each field, various channels can be used to obtain information, including but not limited to patent documents, scientific effect libraries, electromechanical product manuals, etc. Multiple fields can be selected from fields with technical intersections with electromechanical products, fields with rich public data, and fields determined according to actual knowledge requirements, etc. For example, according to the fields with technical intersections with electromechanical product design, fields such as artificial intelligence (optimization of robotic arm structure, joint drive design, etc.) and thermodynamics (the heat dissipation system of electromechanical products involves heat conduction and fluid mechanics) can be determined. According to the fields with rich public data, fields such as patents and scientific effect libraries can be determined. According to actual knowledge requirements, fields such as materials science can be determined because the selection of motor winding materials needs to involve material properties such as electrical conductivity and temperature resistance, and fields such as electronic engineering can be determined because the design of control systems needs to integrate circuit principles and signal processing knowledge.
[0038] After the sample set is constructed, some data in the sample set is labeled to obtain labeled data, and the original sample data that has not been labeled is classified as unlabeled data. Since the labeled data needs to be analyzed and labeled one by one, the number of labeled data can be determined first, and the rest are all used as unlabeled data. For example, if the number of labeled data is determined to be 10,000, then all the data in the sample set except these 10,000 are unlabeled data. Both the labeled data and the unlabeled data can be selected from any field. As an implementation method, the number of labeled data is evenly selected in each field.
[0039] The unlabeled data includes text data. The labeled data includes query statements (such as "how to improve the energy efficiency of motors"), text data (such as patent abstracts, scientific effect descriptions), and the correlation index between the two. The correlation index is an index that quantitatively quantifies the matching degree between the two. For example, a scoring rule of 0 and 1 points or a scoring rule of 0 - 9 points is used to quantify the matching degree, etc. During the labeling process, electromechanical product technicians label the correlation index for each pair of query statements and text data.
[0040] Optionally, before annotating the sample set, data cleaning and preprocessing can also be performed, including noise filtering to remove duplicate text, incorrectly formatted data, and content unrelated to electromechanical products; and / or normalization to unify the text encoding format (such as UTF-8) and eliminate the interference of punctuation marks and case differences on semantic analysis; and / or sentence splitting and word segmentation, using natural language processing tools (such as NLTK, jieba) to split long texts into sentences, providing training samples for subsequent construction of a cross-domain semantic matching model.
[0041] Step S20: Perform incremental pre-training on the pre-trained model according to the unlabeled samples to obtain an intermediate model.
[0042] The pre-trained model is pre-trained based on a general-domain corpus and has basic semantic understanding capabilities. For example, an open-source pre-trained model aspire / acge_text_embedding can be used. The unlabeled samples are input into the pre-trained model for training to obtain an intermediate model. The incremental pre-training process enables the intermediate model to learn the semantic features of electromechanical product knowledge, aligns the intermediate model with multiple fields related to electromechanical products, and improves the cross-domain knowledge retrieval ability of the intermediate model.
[0043] Incremental pre-training is a machine learning strategy that refers to the process of continuing pre-training on an existing model using new datasets or additional data, especially applicable to the fields of deep learning and natural language processing. Incremental pre-training allows the model to gradually accumulate knowledge and adapt to a changing data environment without having to retrain the entire model from scratch.
[0044] During incremental pre-training, a possible hyperparameter setting is: the learning rate is (0.001, 0.0001, 0.00001, 0.000001), the number of training epochs is 20 and an early stopping strategy is set, and the batch size range is (8, 16, 32, 64).
[0045] Step S30: Use a large language model to perform data augmentation on the labeled samples to obtain augmented samples.
[0046] A large language model refers to a deep learning model that can understand and generate human language text and process natural language tasks, such as chat-gpt4, Tongyi Qianwen, Doubao, etc. The choice of the large language model can be adjusted according to the actual situation. The purpose of using the large language model for data augmentation is to increase the diversity of training data to improve the generalization ability of the final model. Data augmentation is performed on the query statements in the labeled samples, and the augmented query statements, together with the original corresponding text data and relevance indicators, form augmented text.
[0047] Different from the methods of synonym replacement and back translation adopted in traditional data augmentation, using a large language model for data augmentation can generate high-quality augmented samples, significantly enhancing the diversity and scale of labeled samples and providing richer training materials for subsequent supervised fine-tuning. Exemplarily, by inputting "Please generate 3 new samples with the same semantic meaning as the following query statement but different expressions, requirements: keep the core intention of the query statement unchanged. The query statement is 'How to improve the motor efficiency'" into the large language model, the augmented samples can be finally obtained.
[0048] Step S40: Supervise and fine-tune the intermediate model based on the labeled samples and augmented samples to obtain a cross-domain semantic matching model.
[0049] In the combined sample set composed of the labeled samples and augmented samples, divide the combined sample set into a training set, a validation set, and a test set according to an appropriate ratio. Input the training set into the intermediate model for supervised fine-tuning to obtain a cross-domain semantic matching model, and use the validation set and the test set for model evaluation.
[0050] In the above solution, training with multiple domain sub-samples related to electromechanical products can explore the potential connections between different domain knowledge, promote the deep integration of knowledge, enable the cross-domain semantic matching model to come into contact with different domain language expressions and knowledge structures, and enhance its adaptability to diverse data. In the incremental pre-training stage, unlabeled samples are used to let the cross-domain semantic matching model learn more extensive semantic features and patterns. When facing cross-domain natural language queries, it can better understand the query intention and match relevant knowledge. In the supervised fine-tuning stage, combined with the labeled samples and augmented samples, the generalization of the cross-domain semantic matching model is further optimized, so that when the model retrieves cross-domain knowledge of electromechanical products, it can more accurately judge the relevance between the query statement and the text data and improve the retrieval accuracy.
[0051] In some embodiments, in step S20, incrementally pre-training the pre-trained model according to the unlabeled samples includes:
[0052] Step S201: Based on the embedding vectors of the unlabeled samples generated by the pre-trained model, use the SimCSE contrastive learning method to construct the first positive sample pairs and the first negative sample pairs.
[0053] As a text embedding model, the pre-trained model can convert text content into vector representations. The pre-trained model processes each unlabeled sample, and through the internal network structure of the model, converts the sample in text form into a numerical embedding vector, and the embedding vector can represent the semantic features of the sample.
[0054] The cross - domain semantic matching model involves knowledge in multiple fields. During incremental pre - training, the SimCSE contrastive learning method is adopted to align the model with the fields related to electromechanical products. The core idea of SimCSE is to input the same unlabeled sample into the pre - trained model twice. During the input process, the dropout random inactivation technique is used to randomly inactivate some neurons in the pre - trained model. In this way, even for the same input sample, due to the random effect of dropout, the embedding vectors output by the model twice will have certain differences, but overall they remain similar because they come from the same sample and their semantic essence is the same. Pairing these two similar but different embedding vectors obtained through dropout forms the first positive sample pair. For example, for the unlabeled sample "Improve the performance of electromechanical equipment by optimizing the motor winding", when input into the model for the first time, the embedding vector u1 is obtained. Under the action of dropout, when input for the second time, the embedding vector u ′ 1, (u1, u ′ 1) forms a first positive sample pair. The construction method of the first positive sample pair simulates different representations of the same semantics under different random conditions, which helps the model learn the core features of the sample semantics without relying on a specific input form.
[0055] For the construction of the first negative sample pair, SimCSE combines the embedding vector of an unlabeled sample pairwise with the embedding vectors of other unlabeled samples within the same training batch. For example, in a batch containing unlabeled samples A, B, and C, the embedding vector of A is u A , the embedding vector of B is u B , and the embedding vector of C is u C , then (u A , u B ) and (u A , u C ) are both first negative sample pairs. In this way, a large number of first negative sample pairs that are semantically unrelated or have low correlation are constructed. During the subsequent training process, the pre - trained model needs to learn how to distinguish between the first positive sample pairs and the first negative sample pairs, continuously optimize its own parameters, so as to enhance the ability to judge semantic similarity and difference, and improve the understanding and matching ability of cross - domain semantics.
[0056] Step S202: Input the similarity of the first positive sample pair and the similarity of the first negative sample pair into the InfoNCE loss function to calculate the loss value.
[0057] After obtaining the first positive sample pair and the first negative sample pair, it is first necessary to calculate the similarity of each pair of sample pairs. For example, the similarity can be measured by distance. In order to make the similarity of the first positive sample pair larger and the similarity of the first negative sample pair smaller, SimCSE uses the InfoNCE loss function for optimization. The similarities of the first positive sample pair and the first negative sample pair are input into the InfoNCE loss function to calculate the loss value. The calculation of the loss value provides a clear direction for the update of the model parameters. By continuously adjusting the model parameters, the embedding vectors generated by the model can better reflect the semantic information of the samples, and the model is aligned with the electromechanical product field.
[0058] Step S203: Optimize the parameters of the pre-trained model by minimizing the loss value of the InfoNCE loss function.
[0059] As an index for measuring the difference between the model prediction result and the ideal result, the loss value of the InfoNCE loss function reflects the gap between the current performance and the target performance of the model. Minimizing the loss value can make the prediction result of the model continuously approach the ideal semantic matching result, thereby realizing the optimization of the model parameters.
[0060] As an implementation, when using an optimization algorithm to complete the adjustment of the pre-trained model parameters, an Adam optimizer can be constructed based on the parameters of the pre-trained model. After testing the optimal hyperparameter configuration of the optimizer on the training set, based on the InfoNCE loss value calculated in step S202, the backpropagation algorithm is used to calculate the gradient of the loss value with respect to the parameters of the pre-trained model. The backpropagation algorithm is based on the chain rule of differentiation and starts from the loss value and propagates backward along the computational graph of the model to calculate the gradient value of each parameter. The Adam optimizer updates the parameters of the pre-trained model according to the calculated gradient value and hyperparameter configuration. Specifically, the optimizer adjusts the parameters in the opposite direction of the gradient, making the InfoNCE loss value gradually decrease.
[0061] In the above solution, using SimCSE contrastive learning to construct the first positive and negative sample pairs based on the unlabeled sample embedding vectors generated by the pre-trained model, and then calculating the loss value through the InfoNCE loss function and optimizing the model parameters, can fully exploit the value of unlabeled data, improve the judgment ability of the cross-domain semantic matching model for semantic similarity, enhance the generalization performance of the model, optimize the training effect, enable the model to better adapt to the cross-domain semantic matching task of electromechanical products, and efficiently and accurately judge the semantic association of texts.
[0062] In some embodiments, step S201 includes:
[0063] Step S2011: In each batch of incremental pre-training, randomly extract a preset number of historical embedding vectors from the experience replay buffer.
[0064] The unlabeled samples are input into the pre-trained model batch by batch, and incremental pre-training is performed on each batch of unlabeled samples. The number of samples in each batch is reasonably set according to the hardware resources and the model training efficiency. For example, in the experimental environment, the batch size can be set to 8, 16, 32, or 64 for testing, and finally a batch size that achieves a balance between the training speed and the effect is determined.
[0065] The experience replay buffer stores the historical embedding vectors generated during the training process. During the training process, whenever the model finishes processing a new batch of unlabeled samples and generates embedding vectors, these embedding vectors are stored in the experience replay buffer. As the training continues, the experience replay buffer will gradually accumulate a large number of historical embedding vectors, which cover the learning results of the model for various text data at different training stages.
[0066] At the beginning of each batch of incremental pre-training, a preset number of historical embedding vectors are randomly sampled from the experience replay buffer. The setting of the preset number can be determined according to the experience of those skilled in the art, and the embodiments of the present application do not make specific limitations in this regard. A suitable preset number enables the model to effectively utilize historical experience while learning new knowledge, reducing the occurrence of catastrophic forgetting. Then, the historical embedding vectors are merged with the new embedding vectors of the current batch, and the merged vectors are jointly used for the incremental pre-training of the current batch. The random sampling method can ensure the diversity of the sampling and avoid the model's dependence on specific types of historical data.
[0067] By combining the historical embedding vectors with the current embedding vectors, the model can take into account both the past learning experience and the current new data during training, thereby better balancing the learning of new knowledge and the retention of old knowledge.
[0068] Step S2012: Construct the first positive sample pair and the first negative sample pair based on the historical embedding vectors and the current embedding vectors.
[0069] Based on the fused historical embedding vectors and the current embedding vectors, the first positive sample pair and the first negative sample pair of the current training batch are constructed.
[0070] Step S203 includes:
[0071] Step S2031: Calculate the gradient of the pre-trained model parameters according to the InfoNCE loss function.
[0072] Step S2032: Perform gradient pruning on the gradient to keep the gradient within a preset range, and the gradient is used to optimize the parameters of the pre-trained model.
[0073] In step S202, the loss value of the InfoNCE loss function has been calculated. The deep learning framework can use the automatic differentiation mechanism to calculate the gradients of the model parameters according to the computational graph of the loss function. To address the problem of gradient explosion that may occur when directly updating the parameters according to the calculated gradients, that is, the gradient value is too large, resulting in unstable model training or even non-convergence, step S2032 uses the method of gradient pruning to limit the calculated gradients to keep them within a preset range. After calculating the gradients, check the magnitude of the gradients. If the norm of the gradients exceeds the preset threshold, scale the gradients so that their norm is equal to the threshold. For example, if the gradient norm threshold is set to 5.0, if the total norm of all parameter gradients exceeds 5.0, all parameter gradients will be scaled by a certain proportion so that the total norm of the scaled gradients is equal to 5.0. It can be understood that the gradient norm threshold can also be set to other values.
[0074] The combined application of experience replay and gradient pruning can combat catastrophic forgetting. Among them, catastrophic forgetting is a key issue in machine learning, especially in the field of deep learning. It refers to the situation where when the model learns a new task, it will significantly forget the knowledge or skills learned previously.
[0075] In the above solution, during incremental pre-training, historical embedding vectors are sampled from the experience replay buffer and combined with the current embedding vectors to construct positive and negative sample pairs, increasing sample diversity, helping the cross-domain semantic matching model to consolidate old knowledge and reduce catastrophic forgetting. At the same time, the gradients calculated by the InfoNCE loss function are pruned to keep the gradients within the preset range, avoiding gradient explosion or disappearance, making the model training more stable, and also reducing catastrophic forgetting, ultimately improving the performance of the pre-trained model in cross-domain semantic matching tasks.
[0076] In some embodiments, the use of the large language model to perform data augmentation on the annotated samples in step S30 includes:
[0077] Sending at least one instruction to the large language model to perform data augmentation on the annotated samples, where the instruction is used to instruct the large language model to output text data with the same semantics as the annotated samples but different expressions.
[0078] There are multiple text contents with the same semantics as the same annotated sample but different expressions. Therefore, the instructions for using the large language model to generate augmented text also include at least one. One instruction is used to generate one augmented text in a specific direction (such as adding, deleting, replacing, etc.). Exemplarily, the following shows three possible instructions:
[0079] Instruction 1:
[0080] "Suppose you are a semantic matching data annotation expert. Now you need to augment the following sample:
[0081] ‘Example sample’
[0082] Requirement: Add 20% redundant information related to the sample without affecting the sample meaning. Just answer the text after adding the redundant information.
[0083] Instruction 2:
[0084] “Suppose you are a semantic matching data annotation expert. Now you need to enhance the following data:
[0085] ‘Example sample’
[0086] Requirement: Delete 20% of the information in the sample without affecting the sample meaning. Just answer the text after deleting the information.
[0087] Instruction 3:
[0088] “Suppose you are a semantic matching data annotation expert. Now you need to enhance the following data:
[0089] ‘Example sample’
[0090] Requirement: Replace 20% of the text in the sample without affecting the sample meaning. Just answer the text after replacement.
[0091] Data augmentation is only carried out for the query statements in the annotated samples. For example, if the query statement A is augmented to obtain three augmented statements A1, A2, and A3 after data augmentation, then A1, A2, and A3 are respectively associated with the text data A corresponding to the query statement A ′ to form three augmented samples (A1, A ′ ), (A2, A ′ ), (A3, A ′ )(the correlation index is not shown). The instruction can be a pre-set template and can be called continuously during data augmentation.
[0092] In the above solution, data augmentation is performed on the annotated samples by sending instructions to the large language model to obtain text data with the same semantics as the annotated samples but different expressions, increasing the diversity of the training data, enabling the cross-domain semantic matching model to learn more diverse language expression patterns during supervised fine-tuning, helping to improve the semantic understanding ability of the model under different expressions, and further enhancing the generalization and accuracy of the model in cross-domain semantic matching tasks.
[0093] In some embodiments, the supervised fine-tuning of the intermediate model according to the annotated samples and the augmented samples in step S40 includes:
[0094] Step S401: Calculate the cosine similarity between the query statements and the corresponding text data in the labeled samples, and the cosine similarity between the query statements and the corresponding text data in the augmented samples.
[0095] Before calculating the cosine similarity, the query statements and the text data need to be converted into vector representations. The cosine similarity measures the similarity between two vectors by calculating the cosine value of the angle between them. For two non-zero vectors A and B, the calculation formula for the cosine similarity cos(θ) is:
[0096]
[0097] For each pair of query statements and text data in the labeled samples and the augmented samples, calculate the cosine similarity of each pair of query statements and text data.
[0098] Step S402: According to the relevance index, scale the cosine similarity between the query statement corresponding to the relevance index and the text data to obtain the scaled cosine similarity.
[0099] Each pair of query statements and text data in the labeled samples and the augmented samples has a relevance index. The higher the relevance index, the higher the semantic matching degree between the corresponding query statement and the text data pair. For example, when using 0 and 1 as the relevance index, 1 indicates that the query statement is highly relevant to the text data; 0 indicates that the query statement has a low relevance to the text data.
[0100] In traditional semantic matching models, the cosine similarity is usually directly used as the similarity measure for sample pairs (in this application, referring to query statements and text data), ignoring the differences in the true matching degrees of different sample pairs, which affects the final matching accuracy.
[0101] During the training process, it may be difficult to judge the semantic similarity of some sample pairs and it is easy to misjudge. Scaling the cosine similarity according to the relevance index can make the scaled cosine similarity more accurately reflect the true semantic matching degree between the query statement and the text data. For example, for a sample pair with a high relevance index but a low cosine similarity calculated by the model, it indicates that the model has difficulty in recognizing this sample pair, so its cosine similarity is reduced through the scaling operation; for a sample pair with a low relevance index but a high cosine similarity calculated by the model, it also indicates that the model has difficulty in recognizing this sample pair, and its cosine similarity is increased through the scaling operation.
[0102] Exemplarily, when the similarity of the predicted positive samples is less than 0.8 after the model iterates 100 times, the scaling factor is 0.8; when the similarity of the predicted negative samples is greater than 0.9, the scaling factor is 1.2; for the remaining samples, the scaling factor is 1, where the scaling factor is used to multiply the cosine similarity predicted by the model.
[0103] Step S403: Supervise and fine-tune the intermediate model based on the scaled cosine similarity.
[0104] The role of the loss function is to measure the degree of difference between the model's prediction results and the actual situation. Through the calculation of the loss function, a loss value that can reflect the current prediction accuracy of the model is obtained. If the model's prediction is accurate, the loss value will be small; if the prediction is inaccurate, the loss value will be large. The loss value of the loss function is calculated using the scaled cosine similarity, and the loss value is used to guide the adjustment and optimization direction of the intermediate model parameters, realizing the supervised fine-tuning of the intermediate model, and finally obtaining a cross-domain semantic matching model.
[0105] In the above solution, by calculating the cosine similarity between the query statement and the text data in the labeled samples and the augmented samples, and scaling the similarity according to the correlation index to obtain the scaled cosine similarity to adjust the weights of the difficult samples during training, and then supervising and fine-tuning the intermediate model based on this, the fine-tuned cross-domain semantic matching model can more accurately capture the semantic matching relationship, strengthen the learning of difficult samples, improve the semantic matching ability of the model on different data, optimize the model performance, and improve the accuracy and efficiency of retrieval.
[0106] In some embodiments, the labeled samples and the augmented samples include second positive sample pairs and second negative sample pairs, where the second positive sample pairs are query statements and text data determined to be highly relevant according to the correlation index and preset rules, and the second negative sample pairs are query statements and text data determined to be lowly relevant according to the correlation index and preset rules.
[0107] The preset rules are used to provide a threshold for the correlation index, and this threshold is the basis for dividing the second positive sample pairs and the second negative sample pairs. For example, when using 0 and 1 as the correlation index, a possible preset rule is: sample pairs with a correlation index of 1 are classified as second positive sample pairs, and sample pairs with a correlation index of 0 are classified as second negative sample pairs.
[0108] The second positive sample pairs are query statement text data pairs with high semantic similarity (hereinafter referred to as sentence pairs), and the second negative sample pairs are sentence pairs with low semantic similarity. During the training process, the model will calculate the cosine similarity between these sentence pairs to determine whether the model's judgment of semantic similarity is accurate. If it is a second positive sample pair, it is expected that the cosine similarity calculated by the model is relatively high; if it is a second negative sample pair, it is expected that the cosine similarity is relatively low. By continuously adjusting the model parameters, the cosine similarity calculated by the model is made to match the true semantic similarity situation of the second positive and negative sample pairs more closely, thereby improving the performance of the model in the semantic matching task.
[0109] In the above solution, by clearly marking the second positive sample pairs and the second negative sample pairs in the samples and the enhanced samples, that is, dividing the query statements and text data with high and low correlations according to the correlation index and the preset rules, it provides a clear positive and negative sample reference for the supervised fine-tuning of the model, enabling the cross-domain semantic matching model to more specifically learn the key features of semantic matching during the training process, strengthening the ability to distinguish between similar and dissimilar text pairs, and effectively improving the accuracy and reliability of the model in cross-domain semantic matching tasks.
[0110] In some embodiments, step S403 includes:
[0111] Determine the loss value according to the scaled cosine similarity and the loss function, and adjust the parameters of the intermediate model according to the loss value to obtain a cross-domain semantic matching model.
[0112] When performing supervised fine-tuning, a possible setting of hyperparameters is: the learning rate is (0.001, 0.0001, 0.00001, 0.000001), the training epoch Epoch is 50 and an early stopping strategy is set, and the batch size batch_size ranges from (8, 16, 32, 64).
[0113] In the above solution, determining the loss value based on the scaled cosine similarity and the loss function, and adjusting the parameters of the intermediate model accordingly to obtain a cross-domain semantic matching model can use the scaled cosine similarity to more accurately reflect the semantic matching degree between the query statement and the text data, make the loss value more accurately measure the difference between the model prediction and the actual situation, guide the model to optimize the parameters, thereby improving the performance of the model in cross-domain semantic matching, enhancing the accuracy and efficiency of semantic matching, and promoting the effective retrieval and application of cross-domain knowledge.
[0114] In some embodiments, the loss function is the CoSENT function with dynamic weights added, and the expression of the CoSENT function is as follows:
[0115]
[0116] Wherein, is the loss value of the CoSENT function, the i, the j, and the sim(i,j) are respectively the query statement, the text data, and the correlation index of the second positive sample pair, the k, the l, and the sim(k,l) are respectively the query statement, the text data, and the correlation index of the second negative sample pair, λ is a hyperparameter, ω kl is the weight factor of the second negative sample pair, ω ij is the weight factor of the second positive sample pair, cos(u k , u l ) is the predicted similarity of the second negative sample pair, cos(u i , uj ) is the predicted similarity of the second positive sample pair.
[0117] The predicted similarity refers to a predicted similarity value output by the intermediate model based on the input data and the current model parameters to predict the similarity relationship between the input sample pairs.
[0118] The difficulty distribution of the second positive and negative sample pairs in the training data is uneven. If all second positive and negative sample pairs are treated equally, the model may over-learn the easily distinguishable sample pairs and ignore the difficult-to-distinguish sample pairs. The weight factors ω kl and ω ij can be dynamically adjusted according to the difficulty of the sample pairs. ω ij controls the importance of the second positive sample pair in the loss function. Increasing ω ij will make the model pay more attention to the similarity of the second positive sample pair, and can make the model more sensitive when distinguishing the second positive sample pair. ω kl controls the importance of the second negative sample pair in the loss function. Increasing ω kl will make the model pay more attention to the similarity of the second negative sample pair, and can make the model more sensitive when distinguishing the second negative sample pair. The purpose of adopting the above method is to assign higher weights to the difficult sample pairs, balance the model's learning of different difficulty sample pairs, accelerate the model convergence, and improve the model accuracy.
[0119] The dynamic weights make the model assign higher weights to the second negative sample pairs with high similarity and the second positive sample pairs with low similarity. As an implementation method of adjusting the dynamic weights, during the training process, when the predicted similarity of the second positive sample pair is less than 0.8 after the model iterates 100 times, then set ω ij to 0.8, and when the predicted similarity of the second negative sample pair is greater than 0.9, then set ω kl to 1.2, and the weights of the remaining sample pairs are set to 1. The reason for adopting the above implementation method is that in the semantic matching task, the ideal output of the model for the second positive sample pair is a higher predicted similarity. When the predicted similarity of the second positive sample pair is less than 0.8, it indicates that the model has difficulty in judging the semantic similarity of such second positive sample pairs, and they may be sample pairs with relatively obscure and complex semantic relationships; similarly, when the predicted similarity of the second negative sample pair is greater than 0.9, it indicates that the model has difficulty in distinguishing the semantic differences of these second negative sample pairs and misjudges them as similar. After setting ω ij and ω kl to 0.8 and 1.2 respectively, the weights of these difficult second positive and negative sample pairs in the loss function calculation are increased, making the model pay more attention to them. Therefore, the model will work harder to learn the features and semantic relationships of these sample pairs in the subsequent training, improve the judgment ability for such sample pairs, and thus improve the overall semantic matching accuracy.
[0120] In the above solution, the CoSENT function with dynamic weights added is used as the loss function, and the scale of the loss function is adjusted by the hyperparameter λ. According to the weight factors of the second positive and negative sample pairs, the importance of the sample pairs in loss calculation is dynamically adjusted, enabling the model to flexibly adjust the learning focus based on the predicted similarity of the sample pairs, assigning higher weights to difficult sample pairs, thereby enhancing the model's ability to distinguish the second positive and negative sample pairs, effectively improving the accuracy of the model in calculating semantic similarity, optimizing the performance in cross-domain semantic matching tasks, and better coping with complex and diverse semantic matching scenarios.
[0121] In some embodiments, after step S40, it further includes:
[0122] Evaluating the cross-domain semantic matching model.
[0123] During the training process, the cross-domain semantic matching model is evaluated on the validation set. By calculating the similarity of the embedding vectors on the validation set and comparing with the true labels, evaluation metrics such as accuracy are obtained. According to the performance of the cross-domain semantic matching model on the validation set, the model with the smallest loss and the highest accuracy is saved and exported in a deployable format.
[0124] This application provides a computer-readable storage medium, including: the computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the methods provided in the above method embodiments.
[0125] This application provides a computer program product, including computer program instructions. When the computer program instructions are read and run by a processor, they execute the methods provided in the above method embodiments.
[0126] Figure 2 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. As Figure 2 shown, the electronic device includes: a processor 201, a memory 202, and a bus 203; wherein, the processor 201 and the memory 202 complete communication with each other through the bus 203; the memory 202 stores program instructions executable by the processor 201, and the processor 201 can execute the methods provided in the above method embodiments by invoking the program instructions.
[0127] The processor 201 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with the ability to process signals. The above-mentioned processor 201 can be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), a network processor (NP), or other conventional processors; it can also be a dedicated processor, including a neural network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Moreover, when there are multiple processors 201, a part of them can be general-purpose processors, and another part can be dedicated processors.
[0128] The memory 202 includes one or more (only one is shown in the figure), which can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The processor 201 and other possible components can access the memory 202 to read and / or write the data therein.
[0129] In particular, one or more computer program instructions can be stored in the memory 202, and the processor 201 can read and run these computer program instructions to implement the weak password scanning behavior recognition method provided by the embodiments of the present application.
[0130] The bus 203 includes one or more (only one is shown in the figure), which can be used to communicate directly or indirectly with other devices for data interaction. The bus 203 may include devices for wired and wireless communication, such as optical fibers, Serial Peripheral Interface (SPI) modules, Inter-Integrated Circuit (I2C), etc., and may also include devices for wireless communication, such as Bluetooth modules, Wi-Fi modules, mobile communication modules (such as 4G, 2G modules), etc.
[0131] It can be understood that Figure 2 The structure shown is only schematic, and the electronic device may also include more or fewer components than those shown Figure 2 in the figure, or have a structure different from that shown Figure 2 in the figure. Figure 2 Each component shown in the figure can be implemented by hardware, software, or a combination thereof. The electronic device may be a physical device, such as a switch, router, server, PC, etc., or may be a virtual device, such as a virtual machine, virtualization container, etc. Moreover, the electronic device is not limited to a single device, and may also be a combination of multiple devices or an integrated environment composed of a large number of devices.
[0132] Figure 3 A cross-domain semantic matching device for electromechanical products provided by an embodiment of the present application includes:
[0133] An acquisition module 300, configured to acquire a sample set, the sample set includes multiple sub-samples related to electromechanical products within a domain, the sub-samples include labeled samples and unlabeled samples, wherein the labeled samples include query statements, text data, and a correlation index between the query statements and the text data.
[0134] An incremental pre-training module 301, configured to perform incremental pre-training on a pre-trained model according to the unlabeled samples to obtain an intermediate model.
[0135] A data augmentation module 302, configured to use a large language model to perform data augmentation on the labeled samples to obtain augmented samples.
[0136] A supervised fine-tuning module 304, configured to perform supervised fine-tuning on the intermediate model according to the labeled samples and the augmented samples to obtain a cross-domain semantic matching model.
[0137] The present application provides a system, on which a cross-domain semantic matching model is deployed; the cross-domain semantic matching model is constructed by using the method steps in any of the above embodiments.
[0138] The above are only embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An electromechanical product cross - domain semantic matching method, characterized in that, Including: Obtain a sample set, where the sample set includes sub-samples related to electromechanical products in multiple fields, and the sub-samples include labeled samples and unlabeled samples. Among them, the labeled samples include query statements, text data, and a correlation index between the query statements and the text data; Perform incremental pre-training on the pre-trained model according to the unlabeled samples to obtain an intermediate model; Use a large language model to perform data augmentation on the labeled samples to obtain augmented samples; Perform supervised fine-tuning on the intermediate model according to the labeled samples and the augmented samples to obtain a cross-domain semantic matching model.
2. The method according to claim 1, wherein The performing incremental pre-training on the pre-trained model according to the unlabeled samples includes: Based on the embedding vectors of the unlabeled samples generated by the pre-trained model, use the SimCSE contrastive learning method to construct the first positive sample pair and the first negative sample pair; Input the similarity of the first positive sample pair and the similarity of the first negative sample pair into the InfoNCE loss function to calculate the loss value; Optimize the parameters of the pre-trained model by minimizing the loss value of the InfoNCE loss function.
3. The method according to claim 2, characterized in that, The using the SimCSE contrastive learning method to construct the first positive sample pair and the first negative sample pair based on the embedding vectors of the unlabeled samples generated by the pre-trained model includes: In each batch of incremental pre-training, randomly extract a preset number of historical embedding vectors from the experience replay buffer; Construct the first positive sample pair and the first negative sample pair based on the historical embedding vectors and the current embedding vectors; The optimizing the parameters of the pre-trained model by minimizing the loss value of the InfoNCE loss function includes: Calculate the gradient of the parameters of the pre-trained model according to the InfoNCE loss function; Perform gradient pruning on the gradient to keep the gradient within a preset range, and the gradient is used to optimize the parameters of the pre-trained model.
4. The method according to claim 1, wherein The using the large language model to perform data augmentation on the labeled samples includes: Send at least one instruction to the large language model to perform data augmentation on the labeled samples, where the instruction is used to instruct the large language model to output text data with the same semantics as the labeled samples but different expression forms.
5. The method according to any one of claims 1-4, characterized in that The performing supervised fine-tuning on the intermediate model according to the labeled samples and the augmented samples includes: Calculate the cosine similarity between the query statement and the corresponding text data in the labeled samples and the cosine similarity between the query statement and the corresponding text data in the augmented samples; According to the correlation index, scale the cosine similarity between the query statement and the text data corresponding to the correlation index to obtain a scaled cosine similarity; Perform supervised fine-tuning on the intermediate model based on the scaled cosine similarity.
6. The method according to claim 5, characterized in that, The labeled samples and the augmented samples include second positive sample pairs and second negative sample pairs, where the second positive sample pairs are query statements and text data determined to be highly relevant according to the relevance index and a preset rule, and the second negative sample pairs are query statements and text data determined to be lowly relevant according to the relevance index and the preset rule.
7. The method according to claim 6, characterized in that, The supervised fine-tuning of the intermediate model based on the scaled cosine similarity includes: Determining a loss value according to the scaled cosine similarity and a loss function, and adjusting the parameters of the intermediate model according to the loss value to obtain the cross-domain semantic matching model.
8. The method according to claim 7, wherein The loss function is a CoSENT function with dynamic weights added, and the expression of the CoSENT function is as follows: Among them, the is the loss value of the CoSENT function, the i, the j, and the sim(i, j) are respectively the query statement, the text data, and the relevance index of the second positive sample pair, the l, the l, and the sim(l, l) are respectively the query statement, the text data, and the relevance index of the second negative sample pair, the λ is a hyperparameter, and the ω kl is the weight factor of the second negative sample pair, and the ω ij is the weight factor of the second positive sample pair, and the cos(u k , u l ) is the predicted similarity of the second negative sample pair, and the cos(u i , u j ) is the predicted similarity of the second positive sample pair.
9. An electronic device, characterized in that, Including: A processor, a memory, and a bus, where the processor and the memory communicate with each other through the bus; The memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1-8 by invoking the program instructions.
10. A computer-readable storage medium, characterized in that, Including: The computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method according to any one of claims 1-8.
11. A computer program product, characterized in that, Including computer program instructions, when the computer program instructions are read and run by a processor, the method according to any one of claims 1-8 is executed.
12. A system, characterized in that, A cross-domain semantic matching model is deployed on the system; the cross-domain semantic matching model is constructed by using the method according to any one of claims 1-8.
Citation Information
Cited By
Vector model fine tuning method and device, equipment, storage medium and program product
CN121502271A
Vector model fine-tuning method, device, equipment, storage medium and program product
CN121502271B