Small sample migration modeling optimization method and system based on comparative learning

By using a few-sample transfer modeling method based on comparative learning, the problem of insufficient stability in cross-domain transfer is solved, and the reliability of cross-domain transfer and rapid model adaptation are achieved, thereby improving the convergence speed and generalization performance of the model.

CN121525535AActive Publication Date: 2026-02-13BAIWEIJINKE (SHANGHAI) INFORMATION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202610056778.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-02-13
Estimated Expiration
2046-01-16

AI Technical Summary

Technical Problem

Existing small-sample modeling methods are difficult to apply stably across devices, acquisition conditions, or operating states, resulting in insufficient quantification of cross-domain transfer stability and unstable convergence during training.

Method used

By employing a few-sample transfer modeling method based on contrastive learning, cross-domain contrastive transfer data is periodically collected and processed with time alignment, noise suppression, anomaly removal, missing data completion, and scale standardization. A fixed-length feature vector set is constructed, a training sample sequence is generated, and a low-dimensional contrastive representation is output through a feature encoding network. The stability of the representation and cross-domain stability are evaluated, a robust sample set is constructed, and the parameters of the feature encoding network are updated.

Benefits of technology

It improves the operability and reliability of cross-domain transfer, avoids the representation drift and semantic discontinuity problems in traditional small-sample transfer, significantly improves the convergence speed and generalization performance of the model, and realizes closed-loop optimization of cross-domain transfer modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121525535A_ABST
    Figure CN121525535A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample migration modeling optimization method and system based on comparative learning, and relates to the technical field of migration data processing, and the method comprises the steps: S1, collecting cross-domain comparative migration data, and carrying out the preprocessing of the collected data; s2, constructing a fixed-length feature vector set, and generating a training sample sequence; s3, constructing a feature coding network, outputting low-dimensional comparison representation, performing quantitative evaluation on representation stability, cross-domain stability and cross-domain complexity of the training sample, and generating a cross-domain robust quantized value; and S4, constructing a robust sample set and a candidate sample set, constructing positive and negative sample pairs according to the low-dimensional contrast representation similarity, evaluating the contrast learning loss of the current training batch, and updating feature coding network parameters based on a loss result. The problems that cross-domain migration stability quantization is insufficient and training process convergence is unstable due to the fact that cross-domain index differences are difficult to process in an existing migration modeling technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of migration data processing, in particular to a small sample migration modeling optimization method and system based on contrast learning. BACKGROUND

[0002] With the continuous popularity of information processing equipment, sensing nodes and computing platforms, various complex engineering systems gradually form a multi-source heterogeneous data environment composed of running state records, process event sequences, terminal interaction signals and environmental change data. These data differ significantly in collection frequency, structure form and noise distribution, making it easy to have problems such as insufficient sample size, inconsistent feature structure and obvious scene shift in the modeling process across devices, regions or tasks. In practical scenarios such as equipment operation monitoring, process control optimization and environmental perception modeling, it is often necessary to uniformly express multi-scene data under limited sample conditions and realize model migration and rapid adaptation across environments.

[0003] For example, the invention with publication number CN120911335A discloses a small sample aerodynamic force modeling method based on multi-task learning; including: constructing a multi-task prediction model, integrating an auxiliary task network and a target task network, and enhancing the feature extraction capability of the encoder through the SE layer and the attention layer; obtaining multi-source aerodynamic data and performing standardization preprocessing; designing a dynamic weight mechanism to adaptively adjust the influence of the auxiliary task prediction on the target task output, realizing effective fusion of low-fidelity and high-fidelity data; performing two-stage training to ensure the efficiency and stability of multi-task learning; verifying the prediction ability of the model under limited sample conditions; through the synergistic effect of multi-task knowledge transfer, feature enhancement and dynamic task integration, combined with the feature optimization characteristics of the SE layer and the attention mechanism, high-precision aerodynamic force prediction is realized under limited sample conditions, providing an efficient and accurate aerodynamic force prediction method for aircraft design.

[0004] For example, the invention with publication number CN119066870A discloses a small sample modeling method for aerodynamic force / torque coefficients based on symbolic regression, including: S1, collecting aerodynamic data of a first aircraft; S2, based on flow state variables, constructing an associated parameter expression of the aerodynamic force / torque coefficients of the aircraft; S3, collecting small sample aerodynamic data of a second aircraft, and optimizing the constant term in the associated parameter expression according to it; S4, using the optimized associated parameter expression to perform polynomial fitting on the small sample aerodynamic data to obtain an aerodynamic force / torque coefficient polynomial expression.

[0005] However, existing small sample modeling methods usually rely on a consistent distribution of data environment, and are difficult to be stably applied in cross-device, cross-acquisition condition or cross-operation state scenarios. In actual engineering systems, there are differences in sampling frequency, non-unified measurement scale, inconsistent noise level and other problems among different acquisition ends, which causes the distribution of features to deviate among different scenarios, resulting in misalignment of representation space and drift of class boundary. Especially when there are not enough samples in new working conditions, it is difficult for the model to build stable cross-domain expression, and it is difficult to meet the needs of complex systems for rapid migration and reliable reasoning.

[0006] Therefore, in view of the above problems, there is an urgent need for a small sample transfer modeling optimization method and system based on contrast learning. SUMMARY

[0007] Technical problems solved In view of the deficiencies of the prior art, the present application provides a small sample transfer modeling optimization method and system based on contrast learning, which solves the problem that the existing transfer modeling technology is difficult to handle cross-domain index differences, resulting in insufficient cross-domain transfer stability quantization and unstable training process convergence.

[0008] Technical scheme To achieve the above purpose, the present application is implemented by the following technical scheme: a small sample transfer modeling optimization method based on contrast learning, comprising: S1, periodically collecting cross-domain contrast transfer data, and performing time alignment, noise suppression, anomaly rejection, missing completion and scale standardization processing on the collected data to generate preprocessed cross-domain contrast transfer data; S2, constructing a fixed-length feature vector set based on the preprocessed cross-domain contrast transfer data according to the event timestamp, and generating a training sample sequence; S3, constructing a feature encoding network based on the fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training sample, generating a cross-domain robust quantization value based on the quantitative evaluation result, and forming a training sample transfer sequence; S4, constructing a robust sample set and a candidate sample set based on the training sample transfer sequence, and constructing a positive and negative sample pair according to the similarity of the low-dimensional contrast representation, evaluating the contrast learning loss of the current training batch, and updating the feature encoding network parameters based on the loss result.

[0009] Further, the specific steps of periodically collecting cross-domain contrast migration data and performing time alignment, noise suppression, abnormality rejection, missing data completion and scale standardization on the collected data are as follows: a fixed-width sliding time window is set as a sampling period, cross-domain contrast migration data is periodically collected, the cross-domain contrast migration data includes terminal interaction duration, single session stay duration, click behavior times, request response time, terminal network downlink rate, event timestamp, adjacent event interval time, event sequence length, platform entry request times, cross-platform forwarding delay, single operation numerical value, period operation numerical value, interface call abnormality times, log writing rate and scene domain number; based on a multi-source timestamp synchronization correction mechanism, the timestamps of different data sources are uniformly corrected, and through a sliding average filtering algorithm, burst fluctuations and measurement noise in the cross-domain contrast migration data are suppressed and smoothed; through an abnormality detection method based on a local outlier factor algorithm, abnormal sampling points are identified and rejected, and based on a K-nearest neighbor interpolation algorithm, local missing data is completed; through a Z-score standardization algorithm, the cross-domain contrast migration data is numerically standardized to unify the dimension scale of different physical quantities.

[0010] Further, the specific steps of constructing a fixed-length feature vector set based on the preprocessed cross-domain contrast migration data according to the event timestamp and generating a training sample sequence are as follows: the preprocessed cross-domain contrast migration data is extracted, all cross-domain contrast migration data is sorted according to the event timestamp, each group of cross-domain contrast migration data is assigned a unique sample index, the cross-domain contrast migration data under the same sample index is spliced into a fixed-length feature vector according to a fixed field order, the corresponding scene domain number is attached as an independent feature dimension in the fixed-length feature vector, and a training sample sequence containing a sample index, a fixed-length feature vector and a scene domain number is formed.

[0011] Further, the specific steps of constructing a feature encoding network based on the fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training samples, generating a cross-domain robust quantization value based on the quantitative evaluation results, and forming a training sample migration sequence are as follows: for the fixed-length feature vectors in the training sample sequence, a multi-layer perceptron neural network algorithm is used to construct a feature encoding network, and the fixed-length feature vectors are sequentially input into each hidden layer neuron to perform linear transformation and nonlinear activation operation, obtaining the corresponding low-dimensional contrast representation; the sample index, fixed-length feature vector, low-dimensional contrast representation and scene domain number are established in a corresponding relationship table according to the same index number; for each training sample, the adjacent sample set with the same scene domain number is retrieved in the corresponding relationship table, the average similarity between the low-dimensional contrast representation of the current training sample and the low-dimensional contrast representations of all adjacent samples is calculated, and the representation stability evaluation value of the current training sample is obtained; for each fixed-length feature vector, the corresponding terminal interaction duration, single session dwell time, click behavior frequency, event sequence length, interface call exception frequency, platform entry request frequency, operation value quantity within a period and single operation value quantity are extracted, and the cross-domain stability evaluation value is calculated; the adjacent event interval time, request response time, terminal network downlink rate, cross-platform forwarding delay and log writing rate are extracted, and the cross-domain complexity evaluation value is calculated; the cross-domain stability evaluation value is multiplied by the representation stability evaluation value, and then divided by the corresponding cross-domain complexity evaluation value, to obtain the cross-domain robust quantization value of the current migration in the corresponding scene domain, and the training samples are sorted according to the cross-domain robust quantization value from large to small to obtain the training sample migration sequence.

[0012] Further, the specific steps of extracting the terminal interaction duration, single session dwell time, click behavior frequency, event sequence length, interface call exception frequency, platform entry request frequency, operation value quantity within a period and single operation value quantity, and comprehensively calculating the cross-domain stability evaluation value are as follows: the terminal interaction duration, single session dwell time and click behavior frequency are summed and added by one to take the natural logarithm, obtaining a basic activity term; the event sequence length is divided by the sum of the event sequence length and a constant one, and then squared to obtain a sequence structure term; the interface call exception frequency is divided by the sum of the platform entry request frequency and a constant one, and the inverse of the obtained ratio is taken as the exponent, and the natural constant e is taken as the base to perform exponential operation, obtaining an exception suppression term; the operation value quantity within a period is divided by the sum of the operation value quantity within a period and the single operation value quantity, obtaining an operation ratio mapping term; the basic activity term, sequence structure term, exception suppression term and operation ratio mapping term are multiplied in turn to obtain the cross-domain stability evaluation value.

[0013] Further, the specific steps of comprehensively calculating the cross-domain complexity evaluation value by extracting the corresponding adjacent event interval time, request response time, terminal network downlink rate, cross-platform forwarding delay and log write rate are as follows: adding one to the adjacent event interval time to obtain a time expansion term; dividing the request response time by the sum of the terminal network downlink rate and a constant one, and then adding the cross-platform forwarding delay and the constant one to obtain a transmission blocking term; adding one to the log write rate to obtain a write disturbance term; and multiplying the time expansion term, the transmission blocking term and the write disturbance term in turn to obtain the cross-domain complexity evaluation value.

[0014] Further, the specific steps of constructing the robust sample set and the candidate sample set based on the training sample migration sequence and constructing the positive and negative sample pairs according to the low-dimensional contrast representation similarity are as follows: based on the training sample migration sequence, the median of the cross-domain robust quantization values of the current batch of training samples is extracted, the training samples with cross-domain robust quantization values not lower than the median are divided into the robust sample set, and the training samples with cross-domain robust quantization values lower than the median are divided into the candidate sample set, and the sample indexes of all training samples in the two sets are recorded respectively; in each training round, N robust samples are extracted from the robust sample set to construct a cross-domain training batch, and the corresponding low-dimensional contrast representation and scene domain number are obtained from the corresponding relationship table through the sample index, and the similarity between the low-dimensional contrast representations of any two robust samples in the training batch is calculated; two sample pairs with the same scene domain number and a low-dimensional contrast representation similarity greater than the similarity upper threshold are recorded as a positive sample pair, and two sample pairs with different scene domain numbers and a low-dimensional contrast representation similarity not greater than the similarity upper threshold are recorded as a negative sample pair.

[0015] Further, the specific steps of evaluating the contrast learning loss of the current training batch are as follows: taking the similarity between the low-dimensional contrast representation of the robust sample and each sample in the positive sample set as the exponent, performing exponential operation with the natural constant e as the base, and then summing all the exponential operation results to obtain a positive sample aggregation term; taking the similarity between the low-dimensional contrast representation of the robust sample and each sample in the positive sample set and the negative sample set as the exponent, performing exponential operation with the natural constant e as the base, and then summing all the exponential operation results to obtain a whole set contrast term; taking the negative natural logarithm of the ratio of the positive sample aggregation term to the whole set contrast term, and multiplying it by the corresponding cross-domain robust quantization value and the representation stability evaluation value to obtain a single sample contrast loss term; after adding all the single sample contrast loss terms, dividing by the sum of the product of the cross-domain robust quantization value and the representation stability evaluation value of all samples in the robust sample set plus one, to obtain the contrast learning loss evaluation value.

[0016] Further, the specific steps of updating the feature encoding network parameters based on the loss result are as follows: based on the contrast learning loss evaluation value, gradient back propagation is performed on the feature encoding network and the encoding network parameters are updated; after the encoding network parameter updating is completed, the candidate samples whose cross-domain robust quantization values are not lower than the quantile threshold and whose low-dimensional contrast representations are not lower than the similarity merging threshold are merged into the robust sample set and the set index is updated; the training batch construction, loss evaluation, parameter updating and sample set maintenance process are cyclically executed, and when the change amplitude of the contrast learning loss evaluation value of the continuous M training rounds is lower than the loss change threshold, the training is terminated and the encoding network parameters are fixed, so as to realize sample migration modeling optimization.

[0017] The second aspect of the application provides a small sample migration modeling optimization system based on contrast learning, comprising: a data acquisition and preprocessing module, a sample index feature generation module, a representation robust score calculation module and a cross-domain robust quantization evaluation module, wherein: the data acquisition and preprocessing module is used for periodically acquiring cross-domain contrast migration data, and performing time alignment, noise suppression, anomaly rejection, missing completion and scale standardization processing on the acquired data to generate preprocessed cross-domain contrast migration data; the sample index feature generation module is used for constructing a fixed-length feature vector set based on the preprocessed cross-domain contrast migration data according to the event timestamp, and generating a training sample sequence; the representation robust score calculation module is used for constructing a feature encoding network based on the fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training sample, generating a cross-domain robust quantization value based on the quantitative evaluation result, and forming a training sample migration sequence; the cross-domain robust quantization evaluation module is used for constructing a robust sample set and a candidate sample set based on the training sample migration sequence, constructing positive and negative sample pairs according to the low-dimensional contrast representation similarity, evaluating the contrast learning loss of the current training batch, and updating the feature encoding network parameters based on the loss result.

[0018] Advantages The application has the following advantages: (1) The small sample migration modeling optimization method and system based on contrast learning, by using a three-dimensional quantitative model of cross-domain stability, cross-domain complexity and representation stability, the multi-source heterogeneous monitoring data of cross-scene domains are uniformly evaluated, so that the originally incomparable cross-domain data have quantifiable robustness features, and the operability and reliability of small sample domain migration are fundamentally improved.

[0019] (2) The small sample transfer modeling optimization method and system based on contrast learning, by simultaneously using the low-dimensional contrast representation output by the encoding network and its similarity structure, a representation stability evaluation system is constructed, so that the samples maintain a highly consistent semantic structure in cross-domain transfer, thereby effectively avoiding common problems such as representation drift and semantic discontinuity in traditional small sample transfer.

[0020] (3) The small sample transfer modeling optimization method and system based on contrast learning, by taking the cross-domain robust quantization value and the representation stability as the weight factors of the loss function, the low-dimensional representation learning can adaptively adjust the optimization direction according to the importance and risk degree of the samples in the cross-domain, thereby significantly improving the convergence speed and generalization performance of the model, and avoiding the convergence difficulty problem of traditional contrast learning in cross-domain tasks.

[0021] (4) The small sample transfer modeling optimization method and system based on contrast learning, by training the dynamic update of the sample transfer sequence and the robust sample set, the model can continuously receive new cross-domain data, recalculate the cross-domain robust quantization value and update the encoding network parameters during running, thereby realizing the closed-loop optimization of cross-domain transfer modeling, and making the model have the ability of continuous learning and adapt to the changing scene domain environment. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 The flowchart of the small sample transfer modeling optimization method based on contrast learning; Figure 2 The structure diagram of the small sample transfer modeling optimization system based on contrast learning; Figure 3 The contrast diagram of cross-domain stability, representation stability and cross-domain complexity evaluation; Figure 4 The column chart of cross-domain robust quantization value in different scene domains. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0024] Please refer to Figures 1-4The embodiment of the application provides a technical scheme: a small sample transfer modeling optimization method based on contrast learning, comprising: S1, periodically collecting cross-domain contrast transfer data, and performing time alignment, noise suppression, abnormality rejection, missing completion and scale standardization processing on the collected data to generate preprocessed cross-domain contrast transfer data; S2, constructing a fixed-length feature vector set based on the preprocessed cross-domain contrast transfer data according to the event timestamp, and generating a training sample sequence; S3, constructing a feature coding network based on the fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training sample, generating a cross-domain robust quantization value based on the quantitative evaluation result, and forming a training sample transfer sequence; S4, constructing a robust sample set and a candidate sample set based on the training sample transfer sequence, constructing a positive and negative sample pair according to the low-dimensional contrast representation similarity, evaluating the contrast learning loss of the current training batch, and updating the feature coding network parameters based on the loss result.

[0025] Specifically, the cross-domain comparison migration data is periodically collected, and time alignment, noise suppression, abnormality rejection, missing completion and scale standardization processing are performed on the collected data. The specific steps of generating the preprocessed cross-domain comparison migration data are as follows: a fixed-width sliding time window is set as a sampling period, the cross-domain comparison migration data is periodically collected, and the cross-domain comparison migration data includes terminal interaction duration, single session stay duration, click behavior times, request response time, terminal network downlink rate, event timestamp, adjacent event interval time, event sequence length, platform entry request times, cross-platform forwarding delay, single operation numerical value, period operation numerical value, interface call abnormal times, log writing rate and scene domain number; wherein: the terminal interaction duration is derived from the time difference between the front-end page loading life cycle and the buried point; the single session stay duration is derived from the time difference between the client session start and end events; the click behavior times are derived from the front-end interaction buried point statistics; the request response time is derived from the time consumption record of the client initiating a request and receiving a response; the terminal network downlink rate is derived from the real-time bandwidth statistics of the operating system network interface; the event timestamp is derived from the event trigger time recorded by the server and the client; the adjacent event interval time is derived from the relative time difference of consecutive events of the same user; the event sequence length is derived from the event number statistics sorted by time; the platform entry request times are derived from the call log count of the service side entry API; the cross-platform forwarding delay is derived from the forwarding time consumption recorded by the link tracking log; the single operation numerical value is directly collected from the numerical input field in the business event processing process; the period operation numerical value is calculated from the cumulative result of all numerical input fields in the same sampling period; the interface call abnormal times are derived from the service side error log count; the log writing rate is derived from the log pipeline writing amount statistics per second; and the scene domain number is derived from the unified mapping of the business configuration table to the scene domain to which the collected data belongs. Based on the multi-source timestamp synchronization correction mechanism, the timestamps of different data sources are uniformly corrected, and the sudden fluctuations and measurement noises in the cross-domain comparison migration data are suppressed and smoothed by the moving average filtering algorithm. In the timestamp synchronization correction process, all event timestamps are offset corrected according to the unified time reference coordinates, so that the data generated by different scene domain nodes, different terminal devices and different collection links can be arranged on the unified time line, avoiding the time misalignment of the cross-domain comparison migration data. The moving average filtering algorithm performs weighted smoothing on adjacent sampling values when executing, so that the terminal interaction duration, single session stay duration, click behavior times, request response time, terminal network downlink rate and other quantitative indicators can be stably output, thereby reducing the disturbance influence caused by link jitter, end-side instantaneous fluctuations and buried point accuracy.The abnormal sampling points are identified and removed by an anomaly detection method based on a local outlier factor algorithm, and the local missing data is completed based on a K-neighbor interpolation algorithm; in the anomaly detection stage, the local density deviation of each cross-domain comparison migration data is calculated, the distortion points in the fields such as log write rate, interface call abnormal number and platform entry request number are identified, and the samples causing deviation of the overall statistical distribution are removed. For the local missing items formed in the removed data segments, the K-neighbor interpolation algorithm is used to find the nearest neighbor feature vector under the same scene domain number, and the interpolation result is generated according to the feature differences such as terminal interaction duration, request response time and adjacent event interval time in the neighborhood, so that the completion content maintains the consistent distribution mode of the scene domain behavior characteristics; wherein K is a positive integer greater than three. The cross-domain comparison migration data is subjected to numerical standardization processing by a Z-score standardization algorithm, and the dimension scale of different physical quantities is unified; in the standardization processing, the overall mean and standard deviation of each feature field are used to generate a standardized numerical value, so that the terminal interaction duration, single session stay duration, click behavior number, request response time, terminal network downlink rate, cross-platform forwarding delay, single operation numerical value, periodic operation numerical value and other data with different dimensions and different orders of magnitude are uniformly expressed in the same scale space, providing a reliable input basis for subsequent construction of fixed-length feature vectors and low-dimensional comparison representation.

[0026] In the embodiment, by unified collection mode and whole-process data preprocessing, an input basis with consistent structure, unified scale and controllable noise is established for the cross-domain comparison migration data, so that the cross-domain comparison migration data with different sources and significant distribution differences is regularized into a feature set with continuous time characteristics and stable statistical attributes before entering the modeling. After this process, the cross-domain comparison migration data shows higher consistency in time dimension, numerical dimension and structure dimension, so that the cross-domain difference can be more truly presented, while avoiding random deviation caused by collection link, terminal behavior and scene fluctuation in the original data, providing a clearer input environment for the subsequent feature coding network to extract migratable representation, and significantly improving the learnability, generalizability and overall robustness of the cross-domain comparison migration modeling.

[0027] Specifically, the specific steps of constructing a fixed-length feature vector set based on the pre-processed cross-domain comparison migration data according to the event timestamp and generating a training sample sequence are as follows: extracting the pre-processed cross-domain comparison migration data, keeping the field order of terminal interaction duration, single session stay time, click behavior times, request response time, terminal network downlink rate, event timestamp, adjacent event interval time, event sequence length, platform entry request times, cross-platform forwarding delay, single operation numerical value, operation numerical value within a period, interface call exception times, log writing rate and scene domain number consistent during the extraction process, ensuring that all sampling period cross-domain comparison migration data enters the processing flow in a unified field structure, sorting all cross-domain comparison migration data according to the event timestamp, and making the page access behavior and link call behavior from different scene domains present a continuous relationship on a unified timeline. Assign a unique sample index to each group of cross-domain comparison migration data, and divide the cross-domain comparison migration data from the same source into the same index according to the continuity of the event timestamp during the assignment process, so that each group of data can maintain the relevant structure of the behavior chain; splice the cross-domain comparison migration data under the same sample index into a fixed-length feature vector according to the fixed field order, and keep the fixed positions of terminal interaction duration, single session stay time, click behavior times, request response time, terminal network downlink rate, adjacent event interval time, event sequence length, platform entry request times, cross-platform forwarding delay, single operation numerical value, operation numerical value within a period, interface call exception times and log writing rate in the feature vector during the splicing process, so that the features of different samples are completely consistent in dimension. And the corresponding scene domain number is attached as an independent feature dimension in the fixed-length feature vector, so that each vector has a clear cross-domain attribute identification while representing user behavior features, providing input basis for subsequent cross-domain difference modeling based on scene domain number, forming a training sample sequence containing sample index, fixed-length feature vector and scene domain number, and enabling the cross-domain comparison migration data to have structured expression capabilities such as indexability, organization and feature coding network training.

[0028] In the present embodiment, the cross-domain comparison migration data is organized based on the unified field structure, continuous time sequence and fixed feature layout, so that the behavior records of different scene domains have consistent time arrangement, stable feature position distribution and clear scene domain attribute expression before entering the modeling link. After the above processing, each data record can present the cross-domain comparison migration data in the form of a fixed-length feature vector, so that the cross-domain behavior difference is clearly expressed in the vector space, effectively improving the organization, indexability and migration of the training sample in the cross-domain environment, and laying a solid structural foundation for the feature coding network to extract stable representation.

[0029] Specifically, the specific steps of constructing a feature encoding network based on a fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training samples, generating a cross-domain robust quantitative value based on the quantitative evaluation results, and forming a training sample migration sequence are as follows: for the fixed-length feature vectors in the training sample sequence, a multi-layer perceptron neural network algorithm is used to construct a feature encoding network, the fixed-length feature vectors are sequentially input into each hidden layer neuron to perform linear transformation and nonlinear activation operation, and the corresponding low-dimensional contrast representation is obtained; when constructing the feature encoding network, the number of layers, the number of neurons and the activation function of the input layer, the hidden layer and the output layer are fixedly configured, so that the fixed-length feature vectors can be compressed from a high-dimensional business space to a contrast representation space through layer-by-layer mapping. In the linear transformation stage, the fixed-length feature vectors are weighted and combined based on the weight matrix, and in the nonlinear activation stage, the activation function is used to strengthen the cross-domain behavior difference, so that the encoding network can extract stable cross-domain representation structure from continuous input, and finally generate low-dimensional contrast representation with compression characteristics in the output layer. The sample index, fixed-length feature vector, low-dimensional contrast representation and scene domain number are established in a corresponding relationship table according to the same index number; when establishing the corresponding relationship table, the terminal interaction duration, single session stay time, click behavior times, request response time, terminal network downlink rate, adjacent event interval time, event sequence length, platform entry request times, cross-platform forwarding delay, single operation numerical value, period operation numerical value, interface call exception times, log writing rate and scene domain number of each sample record can be retrieved with the corresponding low-dimensional contrast representation, so that the subsequent cross-domain difference analysis and adjacent sample retrieval have consistent index structure. For each training sample, retrieve the adjacent sample set with the same scene domain number in the corresponding relationship table, calculate the average similarity between the low-dimensional contrast representation based on the current training sample and the low-dimensional contrast representation of all adjacent samples, and obtain the representation stability evaluation value of the current training sample; in the retrieval process, the samples with the same scene domain number are gathered in the common neighborhood structure according to the vector distance of the low-dimensional contrast representation, so that the representation stability can reflect the behavior consistency within the same scene domain.For each fixed-length feature vector, the corresponding terminal interaction duration, single session stay time, click behavior times, event sequence length, interface call exception times, platform entry request times, operation value amount in period and single operation value amount are extracted, and the cross-domain stability evaluation value is calculated comprehensively; the corresponding adjacent event interval time, request response time, terminal network downlink rate, cross-platform forwarding delay and log writing rate are extracted, and the cross-domain complexity evaluation value is calculated comprehensively; when calculating the cross-domain stability evaluation value, the terminal interaction duration, single session stay time, click behavior times, event sequence length, interface call exception times, platform entry request times, operation value amount in period and single operation value amount are used in turn to construct a stability quantization structure, and when calculating the cross-domain complexity evaluation value, the adjacent event interval time, request response time, terminal network downlink rate, cross-platform forwarding delay and log writing rate are used in turn to construct a complexity quantization structure, so that the cross-domain stability evaluation value and the cross-domain complexity evaluation value can accurately reflect the performance level and difficulty level of the sample in cross-domain migration. The cross-domain stability evaluation value is multiplied by the stability evaluation value, and then divided by the corresponding cross-domain complexity evaluation value to obtain the cross-domain robust quantization value of the current migration of the corresponding scene domain, and the training samples are sorted in descending order of the cross-domain robust quantization value to obtain the migration sequence of the training samples; during the sorting process, the cross-domain robust quantization value of each training sample is kept corresponding to its fixed-length feature vector, low-dimensional contrast representation, scene domain number and sample index, so that the migration sequence can reflect the global sorting structure of the cross-domain behavior performance, and provide executable priority basis for the subsequent training process.

[0030] In the present embodiment, Table 1 is a cross-domain robust quantization value data table, which lists the core quantization indicators of five groups of scene domain samples in the cross-domain migration modeling process, including the cross-domain stability evaluation value, the stability evaluation value, the cross-domain complexity evaluation value, and the cross-domain robust quantization value obtained according to the cross-domain robust quantization value calculation formula. The specific description is as follows: scene domain D1: the cross-domain stability evaluation value is 2.31, the stability evaluation value is 0.87, the cross-domain complexity evaluation value is 1.52, and the cross-domain robust quantization value is 1.32; scene domain D2: the cross-domain stability evaluation value is 1.95, the stability evaluation value is 0.78, the cross-domain complexity evaluation value is 1.36, and the cross-domain robust quantization value is 1.12; scene domain D3: the cross-domain stability evaluation value is 2.68, the stability evaluation value is 0.91, the cross-domain complexity evaluation value is 1.74, and the cross-domain robust quantization value is 1.40; scene domain D4: the cross-domain stability evaluation value is 1.72, the stability evaluation value is 0.65, the cross-domain complexity evaluation value is 1.21, and the cross-domain robust quantization value is 0.92; scene domain D5: the cross-domain stability evaluation value is 3.05, the stability evaluation value is 0.94, the cross-domain complexity evaluation value is 2.03, and the cross-domain robust quantization value is 1.41.

[0031] Table 1 Cross-domain robust quantification value data table

[0032] As Figure 3 shown, the cross-domain stability evaluation value, the representation stability evaluation value and the cross-domain complexity evaluation value from five scene domains are shown to reflect the difference of migration characteristics of different scene domains on three key evaluation indicators. The blue line in the figure corresponds to the cross-domain stability evaluation value, reflecting the stability performance of the scene domain under the multi-dimensional factors of page access, session stay, click behavior, sequence structure, anomaly suppression and operation value quantity. The yellow line in the figure corresponds to the representation stability evaluation value, reflecting the aggregation degree of the low-dimensional contrast representation output by the feature encoding network between the same domain samples, for measuring the intra-domain consistency of the encoding representation. The green line in the figure corresponds to the cross-domain complexity evaluation value, reflecting the complexity level affected by the factors such as event interval, transmission delay, network bandwidth and log write disturbance in the cross-domain migration process. Through the joint visualization of the three types of evaluation values, the differences in stability and complexity between different scene domains can be observed intuitively, providing a basis for subsequent calculation of cross-domain robust quantification value.

[0033] As Figure 4 shown, the cross-domain robust quantification value results of five scene domains after cross-domain migration modeling evaluation are shown. The cross-domain robust quantification value is calculated according to the cross-domain stability evaluation value, the representation stability evaluation value and the cross-domain complexity evaluation value proposed by the present application, wherein the cross-domain stability and the representation stability are positive driving factors, and the cross-domain complexity is a negative inhibiting factor, and the three factors together determine the migratability and robustness of the scene domain samples in the cross-domain migration process. The numerical value marked on each column in the figure is the actual cross-domain robust quantification value of the corresponding scene domain sample, which can intuitively reflect the difference in cross-domain migration adaptability of different scene domains. Figure 4 The cross-domain robust quantification value provides a decision basis for subsequent construction of robust sample set, selection of candidate samples and execution of cross-domain contrast migration optimization training, and is an important reference result for sample selection and training scheduling in the method of the present application.

[0034] In the embodiment, the fixed-length feature vector is encoded layer by layer to generate a low-dimensional contrast representation, so that the cross-domain contrast migration data is uniformly compressed into a structure-stable representation space before entering the migration modeling link, reducing the noise interference and scale difference between high-dimensional features. By establishing a corresponding relationship between the sample index, fixed-length feature vector, low-dimensional contrast representation and scene domain number, the cross-domain behavior data has retrievability and associability in the representation layer. The quantitative evaluation of representation stability, cross-domain stability and cross-domain complexity further enables each sample to be assigned a comparable robustness indicator in a cross-domain environment. The formation of the cross-domain robustness quantitative value explicitly sorts the cross-domain differences. Through the above processing, the training sample migration sequence presents clear structure, clear difference and stable representation, providing a reliable input basis for subsequent priority-based contrast learning and migration optimization, and improving the discriminability and availability of cross-domain migration modeling as a whole.

[0035] Specifically, the specific steps of comprehensively calculating the cross-domain stability evaluation value by extracting the corresponding terminal interaction duration, single session stay duration, click behavior frequency, event sequence length, interface call exception frequency, platform entry request frequency, period operation value quantity and single operation value quantity are as follows: summing the terminal interaction duration, single session stay duration and click behavior frequency, taking the natural logarithm of the sum plus one, strengthening the cumulative expression of cross-domain behavior frequency in the summation stage, and suppressing the deviation caused by abnormally large values in the natural logarithm stage by using the nonlinear compression characteristics of the logarithmic function, to obtain a basic activity item that can represent the strength of cross-domain behavior activity; divide the event sequence length by the sum of the event sequence length and the constant one, and take the square, to maintain the relative proportion of the event sequence structure between different samples in the normalization stage, and highlight the stability difference caused by sequence length changes in the square stage by using the nonlinear amplification ability of the power function, to obtain a sequence structure item that can reflect the coherence of cross-domain behavior; divide the interface call exception frequency by the sum of the platform entry request frequency and the constant one, take the inverse of the obtained ratio as the exponent, and perform exponential operation with the natural constant e as the base, to control the measurement scale of the abnormal proportion in the ratio normalization stage, and make the abnormal proportion greater, the smaller the exponential, in the inverse number exponential stage, to achieve the stability suppression effect of abnormal behavior, to obtain an abnormal suppression item that can measure the strength of abnormal disturbance; divide the period operation value quantity by the sum of the period operation value quantity and the single operation value quantity, to balance the operation quantity structure in the numerical normalization stage, not affected by the sudden increase of single operation, to express the fluctuation of cross-domain operation behavior in a consistent scale, to obtain an operation ratio mapping item that can describe the continuity of operation behavior; multiply the basic activity item, sequence structure item, abnormal suppression item and operation ratio mapping item in turn, to fuse the multi-dimensional characteristics of activity, structure, abnormal suppression and operation behavior continuity in the continuous product stage, to uniformly represent the stability of cross-domain behavior, to obtain the cross-domain stability evaluation value.

[0036] wherein the specific calculation formula of the cross-domain stability evaluation value is: ; In the formula, represents the cross-domain stability evaluation value, represents the terminal interaction duration, represents the single session stay duration, represents the click behavior frequency, represents the event sequence length, represents the interface call exception frequency, represents the platform entry request frequency, represents the operation numerical value amount in a period, represents the single operation numerical value amount.

[0037] In the embodiment, by introducing various normalization operators and nonlinear functions in the construction process of the cross-domain stability evaluation, different cross-domain behavior characteristics are continuously expressed under a unified scale, and a differential reinforcement mechanism for active fluctuations, structural continuity, abnormal disturbances and operation changes is formed under the joint action of exponential operation, logarithmic transformation and square enhancement. The mechanism can establish a stable association mapping between multiple types of cross-domain behavior characteristics, so that the cross-domain stability evaluation value maintains numerical smoothness and discriminant sensitivity when facing strong noise samples, skewed distribution samples and behavior violent fluctuation samples, thereby significantly improving the quantification accuracy and evaluation reliability of the cross-domain behavior stability, and providing more distinguishable basic support for the construction of subsequent cross-domain robust quantization values.

[0038] Specifically, the specific steps of comprehensively calculating the cross-domain complexity evaluation value by extracting the corresponding adjacent event interval time, request response time, terminal network downlink rate, cross-platform forwarding delay and log write rate are as follows: the adjacent event interval time is added by one, the original interval is moderately enlarged in the time dimension, and the linear translation effect of the addition operation is avoided to avoid the numerical folding caused by zero value, to obtain the time expansion term, which can dynamically amplify the trigger density of the cross-domain behavior, thereby strengthening the identification strength of the cross-domain frequency anomaly; the request response time is divided by the sum of the terminal network downlink rate and constant one, and the numerical stability in the low-speed scene is improved by introducing constant one in the denominator, and then the cross-platform forwarding delay and constant one are added, so that the delay in the cross-domain transmission link is overall lifted in the additive coupling, to obtain the transmission blocking term, which can reflect the real blocking strength of the cross-domain link after the superposition of multiple transmission time-consuming factors, and enhance the complexity sensitivity of the high-delay scene under the nonlinear accumulation effect of the addition structure; the log write rate is added by one, so that the write rate is linearly improved in the numerical value, and the complexity collapse caused by the zero write rate is avoided by this way, to obtain the write disturbance term, which can represent the continuous interference caused by log generation in the cross-domain process, so that the write load is amplified in the numerical value; the time expansion term, the transmission blocking term and the write disturbance term are multiplied in turn, so that multiple complexity sources are nonlinearly amplified in the multiplication chain, and the linkage effect between the cross-domain delay, the transmission blocking and the write disturbance is strengthened under the coupling effect of the multiplication structure, to obtain the cross-domain complexity evaluation value.

[0039] wherein the specific calculation formula of the cross-domain complexity evaluation value is: ; In the formula, represents the cross-domain complexity evaluation value, represents the adjacent event interval time, represents the request response time, represents the terminal network downlink rate, represents the cross-platform forwarding delay, represents the log write rate.

[0040] In the embodiment, by introducing time expansion item, transmission delay item and write disturbance item in the cross-domain complexity evaluation process, and sequentially fusing the three types of load characteristics in the nonlinear coupling structure, a joint measurement method is formed which can truly reflect the cross-domain transmission pressure, so that the differences of cross-domain behavior in time dimension, transmission link dimension and write load dimension are unified and mapped into the same complexity scale. The evaluation process achieves synergistic enhancement in numerical stability, sensitivity amplification and correlation influence revelation, so that the cross-domain complexity is no longer dominated by a single factor, but presents more comprehensive dynamic response characteristics, providing a more recognizable complexity basis for the generation of subsequent cross-domain robust quantization values.

[0041] Specifically, the specific steps of constructing the robust sample set and the candidate sample set based on the training sample migration sequence and constructing the positive and negative sample pairs according to the low-dimensional contrast representation similarity are as follows: based on the training sample migration sequence, the median of the cross-domain robust quantization value of the current batch of training samples is extracted, and in the extraction process, the median is used as a segmentation threshold to weaken the influence of extreme high values and extreme low values on the overall distribution, the training samples with the cross-domain robust quantization value not lower than the median are divided into the robust sample set, so that the training samples in the robust sample set as a whole exhibit high cross-domain robustness, and the training samples with the cross-domain robust quantization value lower than the median are divided into the candidate sample set, so that the training samples in the candidate sample set retain potential available samples and avoid directly discarding samples with low robustness, and the sample indexes of all training samples in the two sets are recorded respectively, so as to realize fast association access of the fixed-length feature vector, the low-dimensional contrast representation and the scene domain number through the sample index in the subsequent training rounds; in each training round, N robust samples are extracted from the robust sample set to construct a cross-domain training batch, N is a positive integer greater than three, and in the extraction process, a random sampling strategy is used to ensure that different scene domain numbers have a representative distribution in the training batch, and the corresponding low-dimensional contrast representation and scene domain number are obtained from the corresponding relationship table through the sample index, the similarity between the low-dimensional contrast representations of any two robust samples in the training batch is calculated, and in the calculation process, a similarity measurement method is used to compare the numerical values of the low-dimensional contrast representations, so that the similarity can truly reflect the closeness of different robust samples in the contrast representation space; two sample pairs with a low-dimensional contrast representation similarity greater than the similarity upper threshold and the same scene domain number are recorded as positive sample pairs, and in the labeling of the positive sample pairs, the similarity upper threshold is used to constrain the two samples to have a high closeness relationship in the low-dimensional contrast representation space, and the same scene domain number is used to ensure that the sample pairs come from the same scene domain, so that the positive sample pairs can represent high-similarity behaviors in the same domain; two sample pairs with a low-dimensional contrast representation similarity not greater than the similarity upper threshold and different scene domain numbers are recorded as negative sample pairs, and in the labeling of the negative sample pairs, the condition that the low-dimensional contrast representation similarity is not greater than the similarity upper threshold is used to introduce representation difference, and the condition that the scene domain numbers are different is used to introduce cross-domain difference, so that the negative sample pairs can represent cross-domain low-similarity behaviors, and provide clear positive and negative sample supervision signals for subsequent contrast learning.

[0042] In the embodiment, by introducing the median split strategy of cross-domain robust quantization value in the training sample migration sequence and combining the similarity judgment mode of low-dimensional contrast representation, the structure clear sample contrast relationship can be established in the training process, the construction of positive sample pair and negative sample pair has high cross-domain characteristics, the sample organization inside the training batch is more in line with the structural requirements of cross-domain migration modeling, the subsequent contrast learning can continuously strengthen the same domain similarity and cross-domain difference under the condition of robust sample dominance, the feature encoding network obtains more accurate supervision signal in the training process, the cross-domain representation learning under the condition of small sample has higher consistency and separability, thereby effectively improving the overall robustness of cross-domain migration modeling.

[0043] Specifically, the steps for evaluating the contrastive learning loss of the current training batch are as follows: The similarity between the low-dimensional contrastive representation of the robust sample and each sample in the positive sample set is used as an exponent. An exponential operation is performed with the natural constant e as the base. During the exponential operation, the monotonically enhancing property of the exponential function is used to strengthen and amplify highly similar samples, giving higher weights to positive samples that are closer to the robust samples. Then, all exponential operation results are summed to obtain a positive sample aggregation term, which reflects the degree of concentration of the positive sample set around the robust samples in the low-dimensional contrastive representation space. The low-dimensional contrastive representation of the robust samples... The similarity between each sample in the positive and negative sample sets is used as an exponent, and exponential operations are performed with the natural constant e as the base. During the exponential operation, the sensitivity of the exponential function to different similarity levels is utilized to amplify subtle differences between samples, enabling the positive and negative sample sets to form distinguishable statistical distributions in the numerical space. Then, all exponential operation results are summed to obtain the global comparison term, which comprehensively describes the comparison relationship between robust samples and all samples within the training batch, thus constructing a complete normalized reference benchmark. The ratio of the positive sample aggregation term to the global comparison term is taken as the negative natural logarithm. By leveraging the contraction properties of the logarithmic function to suppress gradient instability caused by excessively large ratios, and using a negative sign to map the similarity structure into a loss form, a larger ratio indicates that the positive sample is closer to the robust sample, thus resulting in a smaller loss. This logarithmic result is then multiplied by the corresponding cross-domain robustness quantization value and representation stability evaluation value. During this multiplication process, the cross-domain robustness quantization value is introduced to adjust the intensity of the overall cross-domain performance of the training samples, and the representation stability evaluation value is introduced to weight the stability of the samples in the low-dimensional contrastive representation space with credibility. This yields a single-sample contrastive loss term that simultaneously reflects cross-domain robustness and representation stability. The overall impact on training intensity is as follows: The cumulative loss value of the training batch is formed by summing all single-sample contrastive loss terms. Then, it is divided by the sum of the cross-domain robust quantization value and the representation stability evaluation value of all samples in the robust sample set plus one. In this normalization process, the balance correction of the differences between samples within the training batch is achieved by the joint accumulation of the cross-domain robust quantization value and the representation stability evaluation value. By adding one, the calculation abnormality caused by the denominator being zero is avoided. Thus, the contrastive learning loss evaluation value is obtained, making the final loss result more stable, controllable and insensitive to the batch sample distribution, providing a reliable gradient basis for subsequent parameter updates.

[0044] The specific formula for calculating the contrastive learning loss evaluation value is as follows: ; In the formula, This represents the comparative learning loss evaluation value. This represents the set of sample indices for the current training batch. Let i represent the set of positive samples of robust sample i. a negative sample set representing robust sample i, a low-dimensional contrast representation similarity between robust samples i and j, a cross-domain robust quantization value of robust sample i, a representation stability evaluation value of robust sample i.

[0045] In this embodiment, the distribution difference of the low-dimensional contrast representation in the positive sample set and the full set sample set is expressed in a combination of exponential operation and logarithmic operation, so that the contrast relationship can form a clear and distinguishable gradient structure in the numerical space. The cross-domain robust quantization value and the representation stability evaluation value are introduced as adaptive weighting factors in the loss construction process, so that the loss function has the dynamic adjustment ability for the cross-domain transfer characteristics. Through this structured design, the training process can strengthen the dominant role of high-quality samples in parameter updating while maintaining numerical stability, making the obtained gradient more reliable and the training convergence behavior more stable, effectively improving the targeted convergence efficiency of cross-domain sample transfer modeling.

[0046] Specifically, the specific steps of updating the feature encoding network parameters based on the loss results are as follows: based on the contrast learning loss evaluation value, the gradient back propagation and the encoding network parameter updating are performed on the feature encoding network. In the gradient back propagation process, the corresponding parameter gradient is generated according to the single sample contrast loss item of each robust sample, and the relative distribution of the low-dimensional contrast representation in the feature space is gradually optimized in the direction of aggregation to the positive sample and separation to the negative sample according to the weight updating mode set in the feature encoding network. After the encoding network parameter updating is completed, the candidate samples whose cross-domain robust quantization value is not less than the quantile threshold value and whose low-dimensional contrast representation similarity with any robust sample is not less than the similarity merging threshold value are selected from the candidate sample set. In the screening process, the low-dimensional contrast representation of the candidate sample is subjected to a similarity check to ensure that the candidate sample has transfer availability in the cross-domain robust quantization value and the low-dimensional contrast representation. The candidate samples that pass the check are merged into the robust sample set and the set index is updated. The training batch construction, loss evaluation, parameter updating and sample set maintenance process are repeatedly performed, and the change amplitude of the contrast learning loss evaluation value of the continuous training rounds is continuously monitored in the loop process. When the change amplitude of the contrast learning loss evaluation value of the continuous M training rounds is less than the loss change threshold, the training is terminated and the encoding network parameters are fixed to realize the optimization of sample transfer modeling; wherein M takes a positive integer value greater than three.

[0047] In the embodiment, by performing gradient back propagation and parameter iterative update driven by contrast learning loss evaluation value in the training process, and combining the cross-domain robust quantization value and the low-dimensional contrast representation similarity to implement dynamic screening and set merging on candidate samples, a continuous self-correction training process with loss convergence trend as the core is realized, so that the feature encoding network can gradually strengthen the positive sample aggregation ability and the negative sample discrimination ability under stable and controllable conditions, and the composition of the robust sample set is optimized step by step with the training process, thereby ensuring that the cross-domain representation quality and convergence reliability of the transfer modeling can still be continuously improved under the condition of small samples, and more general encoding results are provided for subsequent cross-domain transfer reasoning.

[0048] As shown in Figure 2 The second aspect of the present application provides a small sample transfer modeling optimization system based on contrast learning, which comprises a data acquisition and preprocessing module, a sample index feature generation module, a representation robust score calculation module and a cross-domain robust quantization evaluation module. The data acquisition and preprocessing module is used for periodically acquiring cross-domain contrast transfer data, and performing time alignment, noise suppression, abnormal rejection, missing completion and scale standardization processing on the acquired data to generate preprocessed cross-domain contrast transfer data. The sample index feature generation module is used for constructing a fixed-length feature vector set based on the preprocessed cross-domain contrast transfer data according to event timestamps, and generating a training sample sequence. The representation robust score calculation module is used for constructing a feature encoding network based on the fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training sample, generating a cross-domain robust quantization value based on the quantitative evaluation result, and forming a training sample transfer sequence. The cross-domain robust quantization evaluation module is used for constructing a robust sample set and a candidate sample set based on the training sample transfer sequence, constructing positive and negative sample pairs according to the low-dimensional contrast representation similarity, evaluating the contrast learning loss of the current training batch, and updating the feature encoding network parameters based on the loss result.

[0049] In the embodiment, by performing multi-stage processing on the periodically acquired cross-domain contrast transfer data, constructing a fixed-length feature vector with unified structure based on event timestamps, and combining the robustness quantization mechanism of the low-dimensional contrast representation and the sorting structure of the cross-domain robust quantization value to establish the transfer sequence, the robust sample set is dynamically updated in the contrast learning training process, forming a complete closed loop from data acquisition to transfer modeling, so that the generation, evaluation and optimization of cross-domain representation are in a continuous iterative state, the feature encoding network can obtain higher cross-domain robustness under the condition of small samples, and the transfer modeling can show stronger generalization ability when facing scene domain differences, thereby providing systematic technical support for the accuracy improvement and stability enhancement of cross-domain transfer tasks.

[0050] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting; it is not intended to exclude myriad other embodiments of the present application that other presenters can develop. It is also possible, however, that only a single element can be present. It is further noted that such a term as "comprising" is intended to mean that the embodiments include the recited elements, but not excluding other elements. "Consisting essentially of when used herein in relation to a composition, means that the composition includes the recited elements, and can include additional elements, so long as the additional elements do not materially alter the basic and novel characteristics of the claimed composition. "Consisting of" when used herein in relation to a composition, means that the composition includes the recited elements, and no additional elements.

[0051] The preferred embodiments of the application disclosed above are only to help explain the principles of the present application. The preferred embodiments do not describe all the details of the present application, nor limit the present application to only the specific embodiments described. It is apparent that many modifications and variations can be made to the present application based on the content of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is limited only by the claims and their full scope and equivalents.

Claims

1. A small sample transfer modeling optimization method based on contrastive learning, characterized in that, The method comprises the following steps: S1, periodically collecting cross-domain contrast migration data, and performing time alignment, noise suppression, abnormality rejection, missing data completion and scale standardization processing on the collected data to generate preprocessed cross-domain contrast migration data; S2, constructing a fixed-length feature vector set based on the preprocessed cross-domain contrast migration data according to event timestamps, and generating a training sample sequence; S3, constructing a feature encoding network based on the fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training sample, generating a cross-domain robust quantization value based on the quantitative evaluation result, and forming a training sample migration sequence; S4, constructing a robust sample set and a candidate sample set based on the training sample migration sequence, constructing a positive and negative sample pair according to the low-dimensional contrast representation similarity, evaluating the contrast learning loss of the current training batch, and updating the feature encoding network parameters based on the loss result; The specific steps of periodically collecting cross-domain contrast migration data and performing time alignment, noise suppression, abnormality rejection, missing data completion and scale standardization processing on the collected data to generate preprocessed cross-domain contrast migration data are as follows: The fixed-width sliding time window is set as one sampling period, and the cross-domain contrast migration data is periodically collected. The cross-domain contrast migration data includes terminal interaction duration, single session stay time, click behavior times, request response time, terminal network downlink rate, event timestamp, adjacent event interval time, event sequence length, platform entry request times, cross-platform forwarding delay, single operation numerical value, period operation numerical value, interface call abnormal times, log writing rate and scene domain number; The collected cross-domain contrast migration data is uniformly corrected based on a multi-source timestamp synchronization correction mechanism, and the sudden fluctuations and measurement noise in the cross-domain contrast migration data are suppressed and smoothed by a sliding average filtering algorithm. Abnormal sampling points are identified and removed by an abnormality detection method based on a local outlier factor algorithm, and local missing data is completed based on a K-nearest neighbor interpolation algorithm. The cross-domain contrast migration data is numerically standardized by a Z-score standardization algorithm to unify the dimension scale of different physical quantities.

2. The few-shot transfer modeling optimization method based on contrastive learning according to claim 1, characterized in that: The specific steps of constructing a fixed-length feature vector set based on the preprocessed cross-domain contrast migration data according to event timestamps, and generating a training sample sequence are as follows: The preprocessed cross-domain contrast migration data is extracted, all cross-domain contrast migration data is sorted according to event timestamps, each group of cross-domain contrast migration data is assigned a unique sample index, the cross-domain contrast migration data under the same sample index is spliced into a fixed-length feature vector according to a fixed field order, and the corresponding scene domain number is attached as an independent feature dimension in the fixed-length feature vector to form a training sample sequence containing sample index, fixed-length feature vector and scene domain number.

3. The few-shot transfer modeling optimization method based on contrastive learning according to claim 1, characterized in that: The specific steps of constructing a feature encoding network based on a fixed-length feature vector and outputting a low-dimensional contrast representation, quantitatively evaluating the representation stability, cross-domain stability and cross-domain complexity of the training samples, generating a cross-domain robust quantization value based on the quantitative evaluation results, and forming a training sample migration sequence are as follows: For the fixed-length feature vectors in the training sample sequence, a multi-layer perceptron neural network algorithm is used to construct a feature encoding network, and the fixed-length feature vectors are sequentially input into each hidden layer neuron to perform linear transformation and nonlinear activation operation, obtaining the corresponding low-dimensional contrast representation; a corresponding relationship table is established between the sample index, the fixed-length feature vector, the low-dimensional contrast representation and the scene domain number according to the same index number; For each training sample, retrieve the adjacent sample set with the same scene domain number in the corresponding relationship table, calculate the average similarity between the low-dimensional contrast representation of the current training sample and the low-dimensional contrast representations of all adjacent samples based on the current training sample, and obtain the representation stability evaluation value of the current training sample; For each fixed-length feature vector, the corresponding terminal interaction duration, single session dwell time, click behavior frequency, event sequence length, interface call exception frequency, platform entry request frequency, operation value quantity within a period and single operation value quantity are extracted, and the cross-domain stability evaluation value is comprehensively calculated; the adjacent event interval time, request response time, terminal network downlink rate, cross-platform forwarding delay and log writing rate are extracted, and the cross-domain complexity evaluation value is comprehensively calculated; the cross-domain stability evaluation value is multiplied by the representation stability evaluation value, and then divided by the corresponding cross-domain complexity evaluation value, to obtain the cross-domain robust quantization value of the current migration of the corresponding scene domain, and the training samples are sorted according to the cross-domain robust quantization value from large to small to obtain the training sample migration sequence.

4. The few-shot transfer modeling optimization method based on contrastive learning according to claim 3, characterized in that: The specific steps of extracting the corresponding terminal interaction duration, single session dwell time, click behavior frequency, event sequence length, interface call exception frequency, platform entry request frequency, operation value quantity within a period and single operation value quantity, and comprehensively calculating the cross-domain stability evaluation value are as follows: Sum the terminal interaction duration, single session dwell time and click behavior frequency, add one and take the natural logarithm to obtain the basic activity term; divide the event sequence length by the sum of the event sequence length and the constant one, and take the square to obtain the sequence structure term; Divide the interface call exception frequency by the sum of the platform entry request frequency and the constant one, take the inverse of the obtained ratio as the exponent, and perform exponential operation with the natural constant e as the base to obtain the abnormality suppression term; Divide the operation value quantity within a period by the sum of the operation value quantity within a period and the single operation value quantity to obtain the operation ratio mapping term; Multiply the basic activity term, sequence structure term, abnormality suppression term and operation ratio mapping term in turn to obtain the cross-domain stability evaluation value.

5. The few-shot transfer modeling optimization method based on contrastive learning according to claim 3, characterized in that: The specific steps of extracting the corresponding adjacent event interval time, request response time, terminal network downlink rate, cross-platform forwarding delay and log writing rate, and comprehensively calculating the cross-domain complexity evaluation value are as follows: The time expansion term is obtained by adding one to the interval time between adjacent events; the transmission blocking term is obtained by dividing the request response time by the sum of the terminal network downlink rate and a constant one, and adding the cross-platform forwarding delay and the constant one; the write disturbance term is obtained by adding one to the log write rate; and the cross-domain complexity evaluation value is obtained by multiplying the time expansion term, the transmission blocking term and the write disturbance term in sequence.

6. The few-shot transfer modeling optimization method based on contrastive learning according to claim 1, characterized in that: The specific steps of constructing the robust sample set and the candidate sample set based on the training sample migration sequence and constructing the positive and negative sample pairs according to the low-dimensional contrast representation similarity are as follows: Based on the training sample migration sequence, the median of the cross-domain robust quantization value of the current batch of training samples is extracted, the training samples with the cross-domain robust quantization value not lower than the median are divided into a robust sample set, and the training samples with the cross-domain robust quantization value lower than the median are divided into a candidate sample set, and the sample indexes of all training samples in the two sets are recorded respectively; In each training round, N robust samples are extracted from the robust sample set to construct a cross-domain training batch, and the corresponding low-dimensional contrast representation and scene domain number are obtained from the corresponding relationship table through the sample index, and the similarity between the low-dimensional contrast representations of any two robust samples in the training batch is calculated; Two sample pairs with the same scene domain number and the low-dimensional contrast representation similarity greater than the upper similarity threshold are recorded as positive sample pairs, and two sample pairs with different scene domain numbers and the low-dimensional contrast representation similarity not greater than the upper similarity threshold are recorded as negative sample pairs.

7. The few-shot transfer modeling optimization method based on contrastive learning according to claim 1, characterized in that: The specific steps of evaluating the contrast learning loss of the current training batch are as follows: The similarity between the low-dimensional contrast representation of the robust sample and each sample in the positive sample set is taken as an index, and the exponential operation is performed with the natural constant e as the base, and then the sum of all exponential operation results is obtained to obtain the positive sample aggregation term; the similarity between the low-dimensional contrast representation of the robust sample and each sample in the positive sample set and the negative sample set is taken as an index, and the exponential operation is performed with the natural constant e as the base, and then the sum of all exponential operation results is obtained to obtain the whole set contrast term; the negative natural logarithm of the ratio of the positive sample aggregation term to the whole set contrast term is taken, and then multiplied by the corresponding cross-domain robust quantization value and the representation stability evaluation value to obtain the single sample contrast loss term; after adding all single sample contrast loss terms, the sum of the products of the cross-domain robust quantization value and the representation stability evaluation value of all samples in the robust sample set is added to one to obtain the contrast learning loss evaluation value.

8. The few-shot transfer modeling optimization method based on contrastive learning according to claim 1, characterized in that: The specific steps of updating the feature encoding network parameters based on the loss result are as follows: Based on the contrast learning loss evaluation value, gradient backpropagation and encoding network parameter updating are performed on the feature encoding network; after the encoding network parameter updating is completed, the candidate samples with the cross-domain robust quantization value not lower than the quantile threshold and the low-dimensional contrast representation similarity not lower than the similarity merging threshold with any robust sample are selected from the candidate sample set, merged into the robust sample set and the set index is updated; the training batch construction, loss evaluation, parameter updating and sample set maintenance process are repeatedly performed, when the change amplitude of the contrast learning loss evaluation value of the continuous M training rounds is lower than the loss change threshold, the training is terminated and the encoding network parameters are fixed, to realize the sample migration modeling optimization.

9. A small sample transfer modeling optimization system based on contrastive learning, characterized in that: It comprises: The data acquisition and preprocessing module, the sample index feature generation module, the representation robust score calculation module, and the cross-domain robust quantification evaluation module, wherein: The data acquisition and preprocessing module is configured to periodically acquire cross-domain comparison migration data, and perform time alignment, noise suppression, abnormality rejection, missing data completion, and scale standardization processing on the acquired data to generate preprocessed cross-domain comparison migration data. The sample index feature generation module is configured to construct a set of fixed-length feature vectors based on the preprocessed cross-domain comparison migration data according to event timestamps, and generate a training sample sequence. The representation robust score calculation module is configured to construct a feature encoding network based on the fixed-length feature vectors and output low-dimensional comparison representations, quantitatively evaluate the representation stability, cross-domain stability, and cross-domain complexity of the training samples, generate cross-domain robust quantification values based on the quantitative evaluation results, and form a training sample migration sequence. The cross-domain robust quantification evaluation module is configured to construct a robust sample set and a candidate sample set based on the training sample migration sequence, construct positive and negative sample pairs according to the similarity of the low-dimensional comparison representations, evaluate the comparison learning loss of the current training batch, and update the feature encoding network parameters based on the loss result.

Citation Information

Patent Citations

  • Aerodynamic force / torque coefficient small sample modeling method based on sign regression

    CN119066870A

  • Small sample aerodynamic modeling method based on multi-task learning

    CN120911335A

  • Coherence evaluation model training method, coherence evaluation method, coherence evaluation model training device, coherence evaluation method, coherence evaluation device and equipment

    CN116955543A

  • Underwater sound target radiation noise simulation method and system based on generative adversarial network

    CN118821612A

  • Risk control cross-domain risk prediction method and system in combination with transfer learning

    CN120851246A