Business data management method, system and device based on multi-mode loop and medium
By constructing a multimodal feature pool and a dynamic backflow mechanism, the problems of data silos and slow model iteration are solved, and unified representation and value assessment of cross-modal data are realized, enabling real-time adaptive updates and continuous optimization of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, intelligent systems in industries such as finance and healthcare suffer from data silos and one-way model usage, which leads to ineffective cross-modal information collaboration, slow model iteration, inability to achieve lightweight and adaptive updates, stagnation in system intelligence, and inability to build a continuously evolving data flywheel.
By constructing a multimodal feature pool and designing a dynamic backflow mechanism between the temporary storage area and the core storage area, cross-modal feature alignment and fusion are achieved. Predictive labels are generated using historical high-value feature vectors, and high-value sample feature vectors are selected for online incremental updates, forming a closed-loop flow.
It achieves unified representation and value assessment of cross-modal data, reduces reliance on manual annotation, enables the model to adapt to business changes in real time, forms a continuously optimized data flywheel effect, and improves the system's intelligence level.
Smart Images

Figure CN121880822A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and more specifically, to a business data management method, system, device, and medium based on multimodal loops. Background Technology
[0002] In the current intelligent transformation of industries such as finance and healthcare, intelligent systems such as intelligent customer service, OCR recognition, voice quality inspection, automatic claims review, medical image-assisted diagnosis, electronic medical record analysis, and clinical voice input have been widely used. However, these systems generally face the dilemma of "data silos" and "one-way model use," which manifests in the following ways: multi-source heterogeneous data modalities such as voice, text, images, forms (insurance) or images, medical records, and test reports (medical) are fragmented, lacking a unified feature fusion mechanism; massive business data and clinical interaction data are mostly limited to post-event offline analysis, failing to achieve online feedback and continuous optimization, and the data value is not fully explored; model updates heavily rely on periodic retraining and expensive manual annotation, resulting in slow response to business changes or new clinical knowledge and high deployment costs.
[0003] In related technologies, a vertical, isolated, single-task model is typically adopted. For example, voice quality inspection in insurance and image analysis systems in medicine operate independently, and cross-modal information cannot be effectively coordinated. Model iteration is an offline, batch, and slow process, making it difficult to achieve lightweight and adaptive updates using real-time business or diagnostic data. This leads to stagnation in the system's intelligence level and prevents the construction of a continuously evolving "data flywheel." Summary of the Invention
[0004] This disclosure provides at least one business data management method, system, device, and medium based on multimodal loops. By constructing a multimodal feature pool and designing a dynamic backflow mechanism between the temporary storage area and the core storage area, it effectively realizes the closed-loop flow of business data between cross-modal representation, value assessment, and model iteration.
[0005] This disclosure provides a service data management method based on a multimodal loop, including: Obtain multimodal business datasets from business scenarios; wherein, the multimodal business datasets include multimodal business data corresponding to multiple business tasks; Cross-modal feature alignment and fusion processing are performed on the multimodal business data corresponding to each business task in the multimodal business dataset to obtain the feature vector to be evaluated corresponding to each business task, and each feature vector to be evaluated is stored in the temporary storage area of the multimodal feature pool. Based on historical high-value feature vectors stored in the core storage area of the multimodal feature pool, predictive labels are generated for each feature vector to be evaluated in the temporary storage area, and the label confidence of the predicted labels is calculated. Based on the label confidence level, high-value sample feature vectors are selected from the temporary storage area; and the business model used to perform business tasks is updated using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The high-value sample feature vectors are evaluated for value. If the evaluation results meet the predetermined return conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
[0006] This disclosure provides a business data management system based on a multimodal loop, including: A data processing unit is used to acquire a multimodal business dataset from a business scenario; and to perform cross-modal feature alignment and fusion processing on the multimodal business data corresponding to each business task in the multimodal business dataset to obtain the feature vector to be evaluated corresponding to each business task; wherein, the multimodal business dataset includes multimodal business data corresponding to multiple business tasks; The multimodal feature pool includes a temporary storage area and a core storage area; the temporary storage area is used to store the feature vectors to be evaluated; the core storage area is used to store historical high-value feature vectors. The label prediction unit is used to generate predicted labels for each feature vector to be evaluated in the temporary storage area based on the historical high-value feature vectors in the core storage area, and to calculate the label confidence of each predicted label. The model update unit is used to filter out high-value sample feature vectors from the temporary storage area based on the label confidence level; and to update the business model used to perform business tasks using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The sample reflow unit is used to evaluate the value of the high-value sample feature vectors. If the value evaluation result meets the predetermined reflow conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
[0007] This disclosure provides a service data management device based on a multimodal loop, including: The data acquisition module is used to acquire multimodal business datasets from business scenarios; wherein, the multimodal business datasets include multimodal business data corresponding to multiple business tasks; The data processing module is used to perform cross-modal feature alignment and fusion processing on the multimodal business data corresponding to each business task in the multimodal business dataset, to obtain the feature vector to be evaluated corresponding to each business task, and to store each feature vector to be evaluated in the temporary storage area of the multimodal feature pool. The label prediction module is used to generate predicted labels for each feature vector to be evaluated in the temporary storage area based on historical high-value feature vectors stored in the core storage area of the multimodal feature pool, and to calculate the label confidence of the predicted labels. The model update module is used to filter out high-value sample feature vectors from the temporary storage area based on the label confidence level; and to update the business model used to perform business tasks using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The data backflow module is used to evaluate the value of the high-value sample feature vectors. If the value evaluation result meets the predetermined backflow conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
[0008] This disclosure provides a computer device, including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the service data management method based on multimodal loops as described in any of the above possible embodiments is executed.
[0009] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal loop-based service data management method as described in any of the possible embodiments above.
[0010] The business data management method, system, device, and medium based on multimodal loops provided in this disclosure firstly solves the modal fragmentation problem by mapping multi-source heterogeneous data to a unified semantic space through cross-modal feature alignment and fusion technology. Then, based on historical high-value feature vectors, it automatically generates prediction labels with confidence levels for the data to be evaluated, reducing reliance on costly manual annotation. On this basis, it selects high-value samples to perform online incremental updates to the business model, enabling the model to adapt to business changes and new scenarios in real time. Finally, it quantitatively evaluates high-value samples and returns qualified samples to the core storage area, thereby continuously driving model optimization and data quality improvement, forming an increasingly intelligent data flywheel effect.
[0011] In this way, by constructing a multimodal feature pool and designing a dynamic backflow mechanism between the temporary storage area and the core storage area, this disclosure effectively realizes the closed-loop flow of business data between cross-modal representation, value assessment and model iteration.
[0012] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings referenced in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0014] Figure 1 This diagram illustrates an application environment of a service data management method based on a multimodal loop provided in an embodiment of the present disclosure. Figure 2 A flowchart of a service data management method based on a multimodal loop provided in an embodiment of this disclosure is shown; Figure 3 A flowchart of a multimodal service data processing method provided by an embodiment of this disclosure is shown; Figure 4 A flowchart of a business model update method provided by an embodiment of this disclosure is shown; Figure 5 A flowchart of a method for evaluating feature vectors of high-value samples provided in an embodiment of this disclosure is shown; Figure 6 This diagram illustrates the structure of a business data management system based on a multimodal loop, as provided in an embodiment of this disclosure. Figure 7 This diagram illustrates the structure of a service data management device based on a multimodal loop provided in an embodiment of this disclosure. Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0016] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0017] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0018] To facilitate understanding of this embodiment, the execution entity of the service data management method based on multimodal loops provided in this disclosure will first be described in detail. The service data management method based on multimodal loops provided in this embodiment can be applied to applications such as... Figure 1 In this application environment, the client communicates with the server via a network. The client can be a mobile device, user terminal, terminal, handheld device, computing device, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms.
[0019] The service data management method based on multimodal loops provided in this application embodiment will be described in detail below with reference to the accompanying drawings. See also Figure 2 The diagram shows a flowchart of a service data management method based on a multimodal loop provided in this disclosure. The method includes the following steps S201-S205: S201, Obtain multimodal business datasets from business scenarios.
[0020] Understandably, a multimodal business dataset refers to a collection of data in a specific business scenario, consisting of data in multiple different modalities. A modality is the type or representation of the data, such as text, images, audio, video, or tables. These different modalities of business data from various business scenarios can collectively constitute a multimodal business dataset. This dataset contains data corresponding to multiple business tasks. A business task refers to a series of organized activities or operational processes undertaken to achieve a specific business objective, with a clear direction and focus in the business process. For example, in a medical scenario, there might be multiple business tasks such as disease diagnosis and disease prediction; or in a financial scenario, there might be multiple business tasks such as credit assessment and risk warning.
[0021] For example, in the medical field, such as in disease diagnosis, a multimodal business dataset might include patient medical record text information, such as symptom descriptions and past medical history, which directly reflects the patient's basic health status and disease clues; patient medical imaging data, such as X-rays and CT scans, which can visually present the patient's internal structure and lesions; and patient examination audio data, such as cardiac auscultation audio, which can obtain information about the working status of organs such as the heart. Alternatively, in the financial field, taking credit assessment as an example, a multimodal business dataset might include customer text information, such as basic personal information and occupational information, which helps understand the customer's basic background and repayment ability; customer transaction image data, such as scanned copies of transaction vouchers, which can prove the authenticity of the customer's transactions and the flow of funds; and customer voice data, such as voice communication records of customers in telephone customer service, from which information such as the customer's emotions and attitudes can be analyzed to assist in credit assessment.
[0022] In some other embodiments, the business scenario can also be a smart security scenario, a smart transportation scenario, etc., without specific limitations. In a smart security scenario, the corresponding multimodal business data can also include surveillance video data, facial image data, text data recorded by an access control system, etc. The surveillance video data is used to monitor the activities in the area in real time, the facial image data is used for personnel identification, and the text data recorded by the access control system can reflect information such as the time and permissions of personnel entering and exiting. In a smart transportation scenario, the corresponding multimodal business data can also include video data captured by traffic cameras, numerical data collected by vehicle sensors (such as vehicle speed, distance, etc.), and audio data from traffic broadcasts, etc. The traffic camera video data is used to monitor traffic flow and violations, the vehicle sensor numerical data can provide real-time feedback on vehicle operating status, and the traffic broadcast audio data can convey road condition information to drivers.
[0023] S202, perform cross-modal feature alignment and fusion processing on the multimodal business data corresponding to each business task in the multimodal business dataset to obtain the feature vector to be evaluated corresponding to each business task, and store each feature vector to be evaluated in the temporary storage area of the multimodal feature pool.
[0024] Specifically, cross-modal feature alignment refers to matching and mapping data from different modalities at the feature level, enabling comparison and analysis of data from different modalities within the same feature space. Fusion processing integrates features from different modalities to form a comprehensive feature representation, i.e., the feature vector to be evaluated corresponding to the business task. A multimodal feature pool is a data system or database used for centralized storage, indexing, and management of multimodal features. Its temporary storage area, part of its feature storage architecture, is used to temporarily store recently generated feature vectors that have not yet undergone value assessment and confirmation.
[0025] In medical settings, for disease diagnosis tasks, key symptom features in medical records can be aligned with lesion features in medical images. For example, the lung inflammation symptoms described in the text can be matched with the lung inflammation area features shown in CT images. Then, these features from different modalities can be fused into a comprehensive feature vector, which is the feature vector to be evaluated. It not only contains the symptom description information in the text, but also combines the intuitive features of the lesions in the images, which can more comprehensively reflect the patient's condition information.
[0026] In financial scenarios, for credit assessment tasks, credit-related features in customer text information, such as occupational stability, are aligned and integrated with transaction behavior features in transaction images, such as transaction frequency, transaction time, and transaction objects, which can reflect the customer's cash flow and transaction habits, as well as tone features in voice data, such as whether the customer's tone is sincere or impatient in telephone communication. This process yields a feature vector to be evaluated for assessing customer credit.
[0027] In some possible embodiments, multimodal business data may include voice data, text data, image data, and form data corresponding to business tasks; to more effectively achieve cross-modal feature alignment and fusion, and improve the quality and accuracy of the feature vectors to be evaluated, refer to Figure 3 As shown, when performing cross-modal feature alignment and fusion processing on modal service data, the following steps S301~S303 may be included: S301, for the multimodal business data corresponding to each business task, extract the speech semantic features of the speech data; and extract the text semantic features of the text data; and extract the visual and structured joint features of the image data and the form data.
[0028] Specifically, speech semantic features refer to the features extracted from speech data that can express semantic information. For example, after converting speech into text using speech recognition technology, semantic keywords and semantic relationships are further extracted. Text semantic features are extracted directly from text data, including features such as vocabulary, grammar, and semantics. For example, natural language processing technology is used to extract features such as the theme and sentiment of the text. Visual features of image data mainly involve intuitive features such as color, texture, and shape of the image. Structured features of form data refer to features such as the relationships between fields in the form and data format. By jointly extracting the visual and structured features of image data and form data, a more comprehensive set of combined features reflecting the information in the image and form can be obtained.
[0029] Here, pre-trained speech models based on self-supervised learning can be used to extract semantic features from speech data, such as Wav2Vec 2.0 or HuBERT. These models can convert audio waveforms into semantically rich contextual feature vectors without relying on large amounts of transcribed text. Simultaneously, pre-trained language models based on the Transformer architecture, such as BERT or its variant RoBERTa, can be used to extract semantic features from text data. The BERT model, pre-trained on a large corpus, deeply understands the semantic structure of text. When input into text data, the model's output vector representation serves as the text semantic features, accurately reflecting the text's theme, sentiment, and other information. Furthermore, combining visual Transformer models with optical character recognition (OCR) technology can achieve the extraction of joint visual and structured features. The visual Transformer learns global visual features from images, while the OCR engine locates and recognizes text content in images or PDF forms, converting it into structured information that can be further encoded, ultimately fusing them into a unified joint visual and structured feature.
[0030] In some other embodiments, an end-to-end multimodal pre-trained model, such as the VL-BERT model that processes both images and text, or a multimodal base model applicable to images, text, and speech, can be used to jointly complete the above-mentioned joint extraction and alignment tasks of speech, text, image, and form features. The specific implementation method is not limited here.
[0031] S302, map the speech semantic features, text semantic features, and visual and structured joint features of the multimodal business data corresponding to each business task to a shared semantic space.
[0032] Here, the shared semantic space is a unified feature representation space in which features of different modalities have the same dimension and semantic representation. All features are converted into vector forms of the same scale and comparable form, enabling direct comparison and fusion.
[0033] S303, within the shared semantic space, semantic alignment of different modal features from the same business task is performed using a cross-modal contrastive learning method, and the aligned fused feature vector is used as the feature vector to be evaluated for each business task.
[0034] Specifically, cross-modal contrastive learning is a machine learning method that learns feature representations that accurately represent the semantic relationships between different modalities by comparing the similarities and differences between features of different modalities. It maps semantically related features of different modalities to positions close to each other in a shared semantic space, while mapping semantically unrelated features to positions far apart from each other. Then, it integrates the aligned feature representations from multiple modalities from the same business task, for example, by vector concatenation, weighted averaging, or using a lightweight fusion neural network, to generate a unified aligned fusion feature vector, which serves as the feature vector to be evaluated for the business task. By combining the advantages of features of different modalities, it can more accurately reflect the information related to the business task and provide strong support for subsequent business processing.
[0035] In some possible implementations, self-supervised feature enhancement strategies can be used to achieve effective alignment and representation enhancement of features from different modalities in a shared semantic space. The core objective is to enable the model to learn to measure and narrow the semantic distance between different modalities without manual annotation. Specific steps may include the following steps (I) to (II): (I) Construct positive sample pairs for features from different modalities of the same business task, and construct negative sample pairs for features from different modalities of different business tasks; and construct a training sample pair set based on the positive sample pairs and the negative sample pairs. (II) With the goal of minimizing the contrast alignment loss function, a cross-modal feature encoder is trained using the training sample pair set to achieve that the modal features in the positive sample pair are close in distance in the shared semantic space, while the modal features in the negative sample pair are far apart in the shared semantic space.
[0036] Here, training sample pairs for contrastive learning can be constructed based on the correlation between various business tasks. This involves pairwise combinations of modal features extracted from different modalities originating from the same business task to form positive sample pairs. For example, in a medical diagnosis scenario, visual features extracted from a patient's lung CT image and textual descriptions of cough and sputum symptoms in their electronic medical record would form a positive sample pair. Simultaneously, randomly sampled modalities from different business tasks would be paired to form negative sample pairs. For instance, the CT image features of the aforementioned patient could be paired with textual features from another diabetic patient's lab report to form a negative sample pair. Then, a cross-modal feature encoder can be trained by optimizing a contrastive alignment loss function. The optimization objective of this loss function is to maximize the similarity between modal features for each positive sample pair and minimize the similarity between any feature in the positive sample pair and all negative sample features. Specifically, this can be achieved by maximizing the weight of the positive sample pair feature similarity relative to the feature similarity of all negative sample pairs. Thus, by continuously optimizing the contrastive alignment loss function and employing a training mechanism that brings positive pairs closer and pushes negative pairs further away, the model gradually learns to ignore differences in data modalities and focus on their inherent semantic consistency. Ultimately, this achieves semantic alignment of features from different modalities within a shared semantic space. In financial anti-fraud scenarios, this process can be manifested in the model recognizing and bringing together the structured transaction details of a suspicious transaction and the emotional features of customer service communication as semantically similar, while pushing away similar features from a normal transaction. This allows the fused feature vector to more accurately indicate fraud patterns.
[0037] Here, the contrast alignment loss function can be expressed as follows: ; in, Let N be the contrast alignment loss function; N represents the total number of sample pairs in a training batch; and i represents the index of the i-th sample pair. Represented as the anchor feature vector of the i-th sample pair, it can be a feature from any specified modality of a business task, such as text features; Represented as the anchor point feature vector The paired positive sample feature vectors come from a different modality within the same business task as the anchor feature vectors, such as image features paired with the text features mentioned above. Form a positive sample pair; Represented as the feature vector of the k-th negative sample pair, and the feature vector of the current anchor point. Different business tasks can be features of any modality; Represented as a temperature parameter, it is a hyperparameter greater than 0, used to adjust the sharpness of the similarity distribution.
[0038] In some other embodiments, other loss functions may be used as the above-mentioned contrast alignment loss function, such as InfoNCE loss, Triplet loss, etc., which are not specifically limited here.
[0039] In this way, by employing multimodal feature extraction of data from various modalities, mapping heterogeneous features to a shared semantic space, and utilizing cross-modal contrastive learning for precise semantic alignment and fusion, efficient and automated integration and semantic unification of multimodal business data are achieved, breaking down modal barriers and constructing a unified representation with complementary information.
[0040] In some other embodiments, within the shared semantic space, for different modal features from the same business task, the similarity between them can be calculated, and features with high similarity can be aligned to make the features of different modalities more semantically consistent. Then, these aligned features are fused to obtain a fused feature vector, which is the feature vector to be evaluated.
[0041] S203, based on historical high-value feature vectors stored in the core storage area of the multimodal feature pool, generate predicted labels for each feature vector to be evaluated in the temporary storage area, and calculate the label confidence of the predicted labels.
[0042] Understandably, the multimodal feature pool, as a hierarchical feature storage and management center, also includes a core storage area storing high-value feature vectors and their associated business labels that have been accumulated through multiple rounds of historical screening and verification. Historical high-value feature vectors refer to a set of feature vector samples deemed high-quality and representative in past business operations, and their corresponding business labels are usually verified through manual review or a prior high-confidence automatic labeling process. Using these historical high-value feature vectors and their labels, predicted labels can be generated for each feature vector to be evaluated in the temporary storage area, and the label confidence level of the predicted labels can be calculated. Here, the predicted label refers to the most likely business category or result identifier inferred from the content of the feature vector to be evaluated. For example, in an insurance scenario, it might be "vehicle damage claim" or "abnormal health declaration," while in a medical scenario, it might be "pneumonia imaging features" or "diabetic retinopathy." The predicted label confidence level is a quantitative indicator used to measure the reliability and certainty of the predicted label; the higher the value, the more reliable the prediction result.
[0043] For example, when generating predicted labels for each feature vector to be evaluated in the temporary storage area using existing historical high-value feature vectors in the core storage area, and calculating the label confidence of the predicted labels, the following steps (a) to (b) may be included: (a) Using a business task classifier trained based on historical high-value feature vectors in the core storage area of the multimodal feature pool and the business labels corresponding to the historical high-value feature vectors, label prediction is performed on each feature vector to be evaluated in the temporary storage area to determine the predicted label corresponding to each feature vector to be evaluated and the predicted probability distribution corresponding to the predicted label. (b) For each feature vector to be evaluated, the label confidence of the predicted label is calculated based on the predicted probability distribution corresponding to the predicted label of the feature vector to be evaluated and the semantic consistency between different modal business data in the feature vector to be evaluated.
[0044] Specifically, a business task classifier can be trained using historical high-value feature vectors and their corresponding business labels from the core storage area of the multimodal feature pool. This could be a multilayer perceptron classifier, support vector machine, random forest, or a neural network-based classification layer. When a new feature vector to be evaluated is generated, it is input into the classifier, which outputs the probability distribution of the vector belonging to each business category. The category with the highest probability is then used as the predicted label.
[0045] Here, during the training of the business task classifier, historical high-value feature vectors are used as input features, and their verified business labels serve as supervision signals. The classifier parameters are optimized by minimizing prediction loss (such as cross-entropy loss), enabling the classifier to learn to identify patterns related to specific business categories from complex fused features. To address changes in the business environment and concept drift, this training process can be designed to be continuous or periodic. That is, once a certain amount of new high-value feature vectors have accumulated in the core storage area, the business task classifier can be automatically or systematically updated or fine-tuned using this new data. Simultaneously, techniques such as class balancing and hard sample mining can be introduced into the training process to address the uneven distribution of samples across different business categories and improve the classifier's ability to distinguish marginal cases. This ensures that the business task classifier can evolve synchronously with the continuous accumulation of high-quality data assets, maintaining the timeliness and accuracy of its predictions.
[0046] Furthermore, after obtaining the trained business task classifier, each feature vector to be evaluated in the temporary storage area can be used as the input of the classifier. After forward computation, a probability distribution on all possible business categories is output. The category with the highest probability is determined as the predicted label of the vector, and this highest probability value constitutes one of the bases for confidence calculation.
[0047] Here, the calculation of label confidence can rely not only on the probability output by the classifier, i.e., the first confidence, but also on the semantic consistency between the sub-features of each modality that generated the feature vector to be evaluated, i.e., the second confidence. For example, in a car insurance claim task, if the accident location described in the voice highly matches the geographical features shown in the photo, and the vehicle damage recorded in the text matches the damage area in the photo, then the semantic consistency between the modalities is high, and the second confidence is also high. The comprehensive label confidence can be calculated using a weighted formula, for example, the comprehensive label confidence equals the classifier's predicted probability multiplied by the weight coefficient α, plus the modality consistency measure multiplied by the coefficient (1-α).
[0048] In this embodiment, a classifier is trained using historically accumulated high-value feature vectors, and dual-source confidence is calculated by combining the classifier's prediction probability with cross-modal semantic consistency. This enables highly reliable automated business label prediction and quality assessment, thereby significantly reducing reliance on manual annotation and ensuring the quality of data entering subsequent processes.
[0049] S204, based on the label confidence level, select high-value sample feature vectors from the temporary storage area; and use the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors to update the business model used to perform business tasks.
[0050] Specifically, by filtering the label confidence scores of each feature vector to be evaluated in the temporary storage area, high-value sample feature vectors can be selected. High-value sample feature vectors refer to feature vectors with high predicted label confidence scores, high reliability, and potential value. The business model used to perform business tasks is then updated based on these high-value sample feature vectors. In the medical scenario, high-value sample feature vectors that are likely to correspond to real disease conditions are selected based on label confidence scores; for example, feature vectors predicted to be a certain disease with a confidence score higher than 0.8 are selected.
[0051] Here, the business model refers to the algorithmic model used to execute business tasks in actual business scenarios. It processes and analyzes input feature vectors to output business decision results. Continuing the previous example, after selecting feature vectors with a confidence level higher than 0.8, these high-value sample feature vectors and their corresponding predicted labels can be used to update and optimize the disease diagnosis model, enabling it to more accurately identify diseases in subsequent diagnoses. Alternatively, in a financial scenario, feature vectors of customers predicted to have high creditworthiness and high confidence levels can be selected as high-value sample feature vectors. These vectors and their predicted labels can then be used to update the credit assessment model, improving the model's accuracy in assessing customer creditworthiness.
[0052] For example, when filtering high-value sample feature vectors from the temporary storage area based on label confidence, the following steps (1) to (3) may be included: (1) Compare the label confidence level with the preset confidence threshold; (2) Select the feature vectors to be evaluated with a label confidence level not less than the preset confidence threshold as high-value sample feature vectors; (3) Mark the feature vectors to be evaluated that have a label confidence level less than the preset confidence threshold as samples to be reviewed and upload them to the review terminal.
[0053] Here, the preset reliability threshold is a configurable parameter, such as 0.7 or 0.8. Its setting can be dynamically adjusted according to the business's requirements for data quality, the performance of the model at the current stage, and the sufficiency of human review resources. No specific limitation is made here.
[0054] Specifically, when the confidence level of the label of the feature vector to be evaluated is not less than a preset confidence threshold, it indicates that the accuracy of the predicted label of the vector to be evaluated is high enough, and it can be judged as high-reliability data and selected as a high-value sample feature vector. It will be directly output to the subsequent business model update process along with its label as effective data for incremental training, driving the model to quickly adapt to new data patterns. When the confidence level of the label of the feature vector to be evaluated is less than the preset confidence threshold, it indicates that the model judgment corresponding to the vector to be evaluated is ambiguous when predicting the label, there may be contradictions in the information between modalities, or it belongs to a novel case that has not been fully learned. At this time, it can be marked as a sample to be reviewed. These samples to be reviewed, along with their related original multimodal data fragments, predicted labels, and information on the reasons for low confidence (such as a low matching of a certain modality feature), are automatically packaged and uploaded to a dedicated manual review terminal interface. Business experts will conduct final review and annotation. The results after expert review and confirmation or correction will be used as new, verified high-quality data and fed back to the core storage area or the subsequent business model update process.
[0055] Here, by using a clear standard of confidence threshold, an automated preliminary classification of new data is achieved. This ensures that the data flowing into the model training stage has high reliability and usability, thereby guaranteeing the stability and effectiveness of business model updates. Simultaneously, this disclosure also proposes directing samples with high uncertainty and pending value to the manual review stage. This allows valuable manual annotation resources to be concentrated on complex or boundary cases that most require AI intervention, improving overall annotation efficiency and cost-effectiveness.
[0056] In some possible implementations, due to the high real-time requirements, data privacy sensitivity, and regional or group-based differences among different branches or business lines, a cloud-edge collaborative incremental update architecture can be adopted to achieve the goal of quickly adapting to new local sample patterns while maintaining global consistency and stability. Please refer to [link to details]. Figure 4 As shown, when updating the business model used to perform business tasks using high-value sample feature vectors, the following steps S401~S403 may be included: S401, Deploy the main model of the business model in the cloud, and deploy a lightweight incremental learning module based on the main model at the edge.
[0057] Here, the main model of the business model deployed in the cloud data center is a complete model trained on massive amounts of data and possessing strong generalization capabilities. Simultaneously, at the edge, closer to the data source—such as servers in a company's branch offices across different locations, mobile sales devices, or local servers in a hospital—a lightweight incremental learning module based on the main model is deployed. This module is not a complete model copy, but rather a plug-in component containing only a small number of trainable parameters, such as an adapter based on Low-Rank Adaptation (LoRA) or a small neural network adaptation layer. Its core function is to allow the model to fine-tune only a few parameters when absorbing new knowledge, thereby significantly reducing computational and storage overhead and update latency at the edge.
[0058] S402, using the high-value sample feature vector and the predicted label corresponding to the high-value sample feature vector, perform incremental training on the lightweight incremental learning module at the edge end, and update the module parameters of the lightweight incremental learning module.
[0059] Furthermore, once the edge-end business generates and filters out high-value sample feature vectors and their predicted labels, this localized, high-quality data can be used to incrementally train the lightweight incremental learning module at the edge. This training process only updates the parameters of the incremental module itself, while keeping the basic parameters of the main model loaded from the cloud unchanged. In this way, the model can be controlled to quickly absorb local business patterns, such as the impact of specific dialect expressions in a certain region on speech recognition, or the impact of a hospital's unique medical record writing standards on text understanding, thereby enabling the model to adapt instantly and personally.
[0060] S403, the updated module parameters at the edge are periodically synchronized with the main model in the cloud to complete the update of the business model.
[0061] Understandably, to consolidate the beneficial knowledge learned at the edge and feed it back into the global model, while preventing the learning directions of each edge node from diverging, the parameters of the lightweight incremental learning modules deployed at each edge can be periodically synchronized to the cloud after training. The cloud then uses parameter aggregation algorithms, such as weighted averaging or performance-based filtering and merging of incremental module parameters from multiple edge nodes, to effectively integrate these scattered local updates into the main model, ultimately completing the iterative update and enhancement of the global business model. This achieves a two-layer optimization structure that combines the rapid evolution of the micro-model at the edge with the steady improvement of the main model in the cloud.
[0062] S205, the high-value sample feature vector is evaluated for value. If the value evaluation result meets the predetermined return conditions, the high-value sample feature vector is transferred from the temporary storage area to the core storage area of the multimodal feature pool.
[0063] Here, value assessment is a comprehensive evaluation of the importance and effectiveness of high-value sample feature vectors in actual business operations. Evaluation indicators may include the degree of contribution to business decisions, data accuracy, and completeness. Pre-defined backflow conditions are pre-set criteria for determining whether a high-value sample feature vector is worth storing in the core storage area. In a medical scenario, when assessing the value of selected high-value sample feature vectors, their accuracy in disease diagnosis and reliability in disease prediction can be evaluated. If the evaluation results show that the feature vector can significantly improve diagnostic accuracy in disease diagnosis and meet the pre-defined backflow conditions (e.g., an accuracy improvement of more than 10%), it can be transferred from the temporary storage area to the core storage area.
[0064] In other application scenarios, such as in the financial sector, when evaluating the value of high-value sample feature vectors, their performance improvement effect on the credit assessment model can be assessed. If the evaluation results show that the feature vector can reduce the misclassification rate of the credit assessment model to a certain extent and meet the predetermined backflow conditions, such as a reduction in the misclassification rate of more than 5%, it can be transferred to the core storage area.
[0065] For example, to ensure the scientific nature of data repatriation decisions and improve the long-term evolution efficiency of the knowledge base in the core storage area, multi-dimensional quantitative evaluation indicators and automated decision-making processes can be introduced when assessing the value of high-value sample feature vectors. (Refer to...) Figure 5 As shown, when conducting a value assessment to determine whether the predetermined return conditions are met, the following steps S501~S503 may be included: S501, based on the degree of difference between the high-value sample feature vector and the historical high-value feature vector in the core storage area, calculate the sample novelty of the high-value sample feature vector.
[0066] Specifically, the degree of difference can be measured by calculating the average or minimum distance between the feature vector of the high-value sample and all or most similar historical feature vectors in the core storage area. Common calculation methods include the complement of cosine similarity or Euclidean distance. Here, the higher the sample novelty, the lower the degree of repetition between the information pattern carried by the sample and the existing knowledge base. It has higher potential value for expanding the cognitive boundary of the model and preventing the model from overfitting to common patterns. For example, in the medical scenario, a tumor feature vector that exhibits a rare combination of image texture and special biomarkers will have a large difference from the feature vectors of historical typical cases, and therefore has high novelty.
[0067] S502, based on the performance changes of the business model in the validation set and / or online business after the model is updated by introducing the high-value sample feature vector, calculate the business impact of the high-value sample feature vector.
[0068] Here, business impact is a post-hoc or projected utility evaluation metric. In practice, it can be implemented through offline or online A / B testing: a small batch of data containing the sample is used for model fine-tuning, and then, keeping other conditions constant, the improvement in key business metrics is evaluated on an independent validation set, such as an improvement in accuracy for an insurance claims review model, an improvement in recall for an anti-fraud model, or an improvement in AUC for a disease diagnosis model. The greater the improvement, the higher the business impact score for that sample. This ensures that the feedback decision is strongly correlated with the final business effect, and can filter out samples that truly contribute positively to the model's capabilities.
[0069] S503, calculate the value assessment result of the feature vector of the high-value sample based on the label confidence, the sample novelty, and the business impact.
[0070] Specifically, after calculating the sample novelty and business impact corresponding to the feature vectors of high-value samples, the value assessment result of the feature vectors of high-value samples can be calculated by combining their own label confidence. Among them, label confidence ensures the reliability of the data, sample novelty encourages the diversity of the knowledge base, and business impact anchors the business benefit orientation of decision-making.
[0071] For example, during the calculation process, configurable weight coefficients can be assigned to each of the three dimensions, and a final comprehensive value score can be calculated using a weighted aggregation function (e.g., linear weighted summation). Only when the comprehensive value score exceeds a preset reflux threshold can the high-value sample feature vector be determined to meet the predetermined reflux conditions, thereby formally migrating it from the temporary storage area to the core storage area of the multimodal feature pool, becoming a new historical high-value feature vector in subsequent iterations.
[0072] In this embodiment of the disclosure, the above-mentioned automated data optimization and accumulation method ensures that every piece of data flowing back to the core knowledge base is high-quality data that has been tested by multiple standards, thereby driving the entire "data flywheel system" to continuously and efficiently evolve towards higher performance, wider coverage and stronger adaptability.
[0073] It should be noted that the business data management method based on multimodal loops proposed in this disclosure is designed with modality independence and business universality. The closed-loop logic constructed by this method for data acquisition, feature alignment, intelligent annotation, model updating, and value feedback can be widely applied to any business scenario involving multi-source heterogeneous data and requiring continuous intelligence, such as intelligent underwriting and claims automation in financial insurance, assisted diagnosis and patient management in the medical and health field, and credit risk assessment and anti-fraud monitoring in financial services. In specific applications, it is only necessary to connect the multimodal data such as voice, text, images, and forms generated in the target business scenario to the system's data acquisition layer, and clarify the specific business tasks and objectives in the scenario. This will automatically initiate cross-modal fusion and self-evolution flywheel, realizing the automatic accumulation of business knowledge and continuous optimization of model capabilities. The specific business type, data modality combination, and model architecture can be adjusted according to actual needs, and are not specifically limited here.
[0074] The business data management method, system, device, and medium based on multimodal loops provided in this disclosure effectively realize the closed-loop flow of business data between cross-modal representation, value assessment, and model iteration by constructing a multimodal feature pool and designing a dynamic backflow mechanism between the temporary storage area and the core storage area.
[0075] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0076] Based on the same inventive concept, this disclosure also provides a multimodal loop-based business data management system corresponding to the multimodal loop-based business data management method. Since the principle of the system in this disclosure for solving the problem is similar to the multimodal loop-based business data management method described above, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.
[0077] Reference Figure 6 The diagram shown is a schematic of a business data management system based on a multimodal loop provided in an embodiment of this disclosure. The system includes: A data processing unit is used to acquire a multimodal business dataset from a business scenario; and to perform cross-modal feature alignment and fusion processing on the multimodal business data corresponding to each business task in the multimodal business dataset to obtain the feature vector to be evaluated corresponding to each business task; wherein, the multimodal business dataset includes multimodal business data corresponding to multiple business tasks; The multimodal feature pool includes a temporary storage area and a core storage area; the temporary storage area is used to store the feature vectors to be evaluated; the core storage area is used to store historical high-value feature vectors. The label prediction unit is used to generate predicted labels for each feature vector to be evaluated in the temporary storage area based on the historical high-value feature vectors in the core storage area, and to calculate the label confidence of each predicted label. The model update unit is used to filter out high-value sample feature vectors from the temporary storage area based on the label confidence level; and to update the business model used to perform business tasks using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The sample reflow unit is used to evaluate the value of the high-value sample feature vectors. If the value evaluation result meets the predetermined reflow conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
[0078] In some possible embodiments, the multimodal service data includes voice data, text data, image data, and form data corresponding to the service task; the data processing unit is specifically used for: For each business task, extract the speech semantic features of the speech data; extract the text semantic features of the text data; and extract the combined visual and structured features of the image data and the form data. The speech semantic features, text semantic features, and visual and structured joint features of the multimodal business data corresponding to each business task are mapped to a shared semantic space. Within the shared semantic space, cross-modal contrastive learning is used to semantically align features from different modalities of the same business task, and the aligned fused feature vector is used as the feature vector to be evaluated for each business task.
[0079] In some possible embodiments, the data processing unit is specifically used for: Features from different modalities of the same business task are paired to form positive sample pairs, and features from different modalities of different business tasks are paired to form negative sample pairs; and a training sample pair set is constructed based on the positive sample pairs and the negative sample pairs. With the goal of minimizing the contrastive alignment loss function, a cross-modal feature encoder is trained using the training sample pair set to achieve that the modal features in the positive sample pair are close in distance in the shared semantic space, while the modal features in the negative sample pair are far apart in the shared semantic space. The contrastive alignment loss function is used to measure the weight of the similarity between features of positive sample pairs relative to the similarity between features of all negative sample pairs, and semantic alignment of features of different modalities is achieved in the shared semantic space by optimizing the contrastive alignment loss function.
[0080] In some possible embodiments, the label prediction unit is specifically used for: A business task classifier trained based on historical high-value feature vectors in the core storage area of the multimodal feature pool and the business labels corresponding to the historical high-value feature vectors is used to predict the labels of each feature vector to be evaluated in the temporary storage area, thereby determining the predicted labels and the predicted probability distributions of each feature vector to be evaluated. For each feature vector to be evaluated, the label confidence of the predicted label is calculated based on the predicted probability distribution corresponding to the predicted label of the feature vector to be evaluated and the semantic consistency between different modal business data in the feature vector to be evaluated.
[0081] In some possible embodiments, the sample reflux unit is specifically used for: The sample novelty of the high-value sample feature vector is calculated based on the degree of difference between the high-value sample feature vector and the historical high-value feature vector in the core storage area. Based on the performance changes of the business model in the validation set and / or online business after updating the model by introducing the high-value sample feature vector, the business impact of the high-value sample feature vector is calculated. The value assessment result of the feature vector of the high-value sample is calculated based on the label confidence, the sample novelty, and the business impact.
[0082] In some possible embodiments, the model update unit is specifically used for: Deploy the main model of the business model in the cloud, and deploy a lightweight incremental learning module based on the main model at the edge. Using the feature vector of the high-value sample and the predicted label corresponding to the feature vector of the high-value sample, the lightweight incremental learning module is incrementally trained at the edge end to update the module parameters of the lightweight incremental learning module; The updated module parameters at the edge are periodically synchronized with the main model in the cloud to complete the update of the business model.
[0083] Based on the same inventive concept, this disclosure also provides a service data management device based on multimodal loops, which corresponds to the service data management method based on multimodal loops. Since the principle of the device in this disclosure is similar to the service data management method based on multimodal loops described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0084] Reference Figure 7 The diagram shown is a schematic of a service data management device 700 based on a multimodal loop provided in this disclosure embodiment. The device includes: The data acquisition module 701 is used to acquire a multimodal business dataset from a business scenario; wherein, the multimodal business dataset includes multimodal business data corresponding to multiple business tasks; The data processing module 702 is used to perform cross-modal feature alignment and fusion processing on the multimodal business data corresponding to each business task in the multimodal business dataset, to obtain the feature vector to be evaluated corresponding to each business task, and to store each feature vector to be evaluated in the temporary storage area of the multimodal feature pool. The label prediction module 703 is used to generate predicted labels for each feature vector to be evaluated in the temporary storage area based on historical high-value feature vectors stored in the core storage area of the multimodal feature pool, and to calculate the label confidence of the predicted labels. The model update module 704 is used to filter out high-value sample feature vectors from the temporary storage area based on the label confidence; and to update the business model used to perform business tasks using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The data backflow module 705 is used to evaluate the value of the high-value sample feature vectors. If the value evaluation result meets the predetermined backflow conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
[0085] In some possible embodiments, the multimodal service data includes voice data, text data, image data, and form data corresponding to the service task; the data processing module 702 is specifically used for: For each business task, extract the speech semantic features of the speech data; extract the text semantic features of the text data; and extract the combined visual and structured features of the image data and the form data. The speech semantic features, text semantic features, and visual and structured joint features of the multimodal business data corresponding to each business task are mapped to a shared semantic space. Within the shared semantic space, cross-modal contrastive learning is used to semantically align features from different modalities of the same business task, and the aligned fused feature vector is used as the feature vector to be evaluated for each business task.
[0086] In some possible embodiments, the data processing module 702 is specifically used for: Features from different modalities of the same business task are paired to form positive sample pairs, and features from different modalities of different business tasks are paired to form negative sample pairs; and a training sample pair set is constructed based on the positive sample pairs and the negative sample pairs. With the goal of minimizing the contrastive alignment loss function, a cross-modal feature encoder is trained using the training sample pair set to achieve that the modal features in the positive sample pair are close in distance in the shared semantic space, while the modal features in the negative sample pair are far apart in the shared semantic space. The contrastive alignment loss function is used to measure the weight of the similarity between features of positive sample pairs relative to the similarity between features of all negative sample pairs, and semantic alignment of features of different modalities is achieved in the shared semantic space by optimizing the contrastive alignment loss function.
[0087] In some possible embodiments, the label prediction module 703 is specifically used for: A business task classifier trained based on historical high-value feature vectors in the core storage area of the multimodal feature pool and the business labels corresponding to the historical high-value feature vectors is used to predict the labels of each feature vector to be evaluated in the temporary storage area, thereby determining the predicted labels and the predicted probability distributions of each feature vector to be evaluated. For each feature vector to be evaluated, the label confidence of the predicted label is calculated based on the predicted probability distribution corresponding to the predicted label of the feature vector to be evaluated and the semantic consistency between different modal business data in the feature vector to be evaluated.
[0088] In some possible embodiments, the data return module 705 is specifically used for: The sample novelty of the high-value sample feature vector is calculated based on the degree of difference between the high-value sample feature vector and the historical high-value feature vector in the core storage area. Based on the performance changes of the business model in the validation set and / or online business after updating the model by introducing the high-value sample feature vector, the business impact of the high-value sample feature vector is calculated. The value assessment result of the feature vector of the high-value sample is calculated based on the label confidence, the sample novelty, and the business impact.
[0089] In some possible embodiments, the model update module 704 is specifically used for: Deploy the main model of the business model in the cloud, and deploy a lightweight incremental learning module based on the main model at the edge. Using the feature vector of the high-value sample and the predicted label corresponding to the feature vector of the high-value sample, the lightweight incremental learning module is incrementally trained at the edge end to update the module parameters of the lightweight incremental learning module; The updated module parameters at the edge are periodically synchronized with the main model in the cloud to complete the update of the business model.
[0090] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 8 The diagram shows the structure of a computer device 800 provided in this embodiment of the present disclosure, including a processor 801, a memory 802, and a bus 803. The memory 802 stores execution instructions and includes a main memory 8021 and an external memory 8022. The main memory 8021, also called internal memory, is used to temporarily store computational data in the processor 801 and data exchanged with external memory 8022 such as a hard disk. The processor 801 exchanges data with the external memory 8022 through the main memory 8021.
[0091] In this embodiment, the memory 802 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 801. That is, when the computer device 800 is running, the processor 801 communicates with the memory 802 through the bus 803, so that the processor 801 executes the application code stored in the memory 802, and then executes the method described in any of the foregoing embodiments.
[0092] The memory 802 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0093] Processor 801 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0094] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 800. In other embodiments of this application, the computer device 800 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0095] This disclosure also provides a computer-readable storage medium storing a computer program. When a processor executes the computer program, it performs the steps of the multimodal loop-based service data management method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0096] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the business data management method based on multimodal loops described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0097] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0100] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0101] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0102] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A service data management method based on multimodal loops, characterized in that, include: Obtain multimodal business datasets from business scenarios; wherein, the multimodal business datasets include multimodal business data corresponding to multiple business tasks; Cross-modal feature alignment and fusion processing are performed on the multimodal business data corresponding to each business task in the multimodal business dataset to obtain the feature vector to be evaluated corresponding to each business task, and each feature vector to be evaluated is stored in the temporary storage area of the multimodal feature pool. Based on historical high-value feature vectors stored in the core storage area of the multimodal feature pool, predictive labels are generated for each feature vector to be evaluated in the temporary storage area, and the label confidence of the predicted labels is calculated. Based on the label confidence level, high-value sample feature vectors are selected from the temporary storage area; and the business model used to perform business tasks is updated using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The high-value sample feature vectors are evaluated for value. If the evaluation results meet the predetermined return conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
2. The method according to claim 1, characterized in that, The multimodal business data includes voice data, text data, image data, and form data corresponding to the business tasks; the cross-modal feature alignment and fusion processing of the multimodal business data corresponding to each business task in the multimodal business dataset includes: For each business task, extract the speech semantic features of the speech data; extract the text semantic features of the text data; and extract the combined visual and structured features of the image data and the form data. The speech semantic features, text semantic features, and visual and structured joint features of the multimodal business data corresponding to each business task are mapped to a shared semantic space. Within the shared semantic space, cross-modal contrastive learning is used to semantically align features from different modalities of the same business task, and the aligned fused feature vector is used as the feature vector to be evaluated for each business task.
3. The method according to claim 2, characterized in that, The semantic alignment of different modal features from the same business task using a cross-modal contrastive learning method includes: Features from different modalities of the same business task are paired to form positive sample pairs, and features from different modalities of different business tasks are paired to form negative sample pairs; and a training sample pair set is constructed based on the positive sample pairs and the negative sample pairs. With the goal of minimizing the contrastive alignment loss function, a cross-modal feature encoder is trained using the training sample pair set to achieve that the modal features in the positive sample pair are close in distance in the shared semantic space, while the modal features in the negative sample pair are far apart in the shared semantic space. The contrastive alignment loss function is used to measure the weight of the similarity between features of positive sample pairs relative to the similarity between features of all negative sample pairs, and semantic alignment of features of different modalities is achieved in the shared semantic space by optimizing the contrastive alignment loss function.
4. The method according to claim 1 or 3, characterized in that, The step of generating predicted labels for each feature vector to be evaluated in the temporary storage region and calculating the label confidence of the predicted labels includes: A business task classifier trained based on historical high-value feature vectors in the core storage area of the multimodal feature pool and the business labels corresponding to the historical high-value feature vectors is used to predict the labels of each feature vector to be evaluated in the temporary storage area, thereby determining the predicted labels and the predicted probability distributions of each feature vector to be evaluated. For each feature vector to be evaluated, the label confidence of the predicted label is calculated based on the predicted probability distribution corresponding to the predicted label of the feature vector to be evaluated and the semantic consistency between different modal business data in the feature vector to be evaluated.
5. The method according to claim 4, characterized in that, The value assessment of the feature vectors of the high-value samples includes: The sample novelty of the high-value sample feature vector is calculated based on the degree of difference between the high-value sample feature vector and the historical high-value feature vector in the core storage area. Based on the performance changes of the business model in the validation set and / or online business after updating the model by introducing the high-value sample feature vector, the business impact of the high-value sample feature vector is calculated. The value assessment result of the feature vector of the high-value sample is calculated based on the label confidence, the sample novelty, and the business impact.
6. The method according to claim 1, characterized in that, The step of updating the business model for performing business tasks using the feature vectors of the high-value samples and the predicted labels corresponding to the feature vectors of the high-value samples includes: Deploy the main model of the business model in the cloud, and deploy a lightweight incremental learning module based on the main model at the edge. Using the feature vector of the high-value sample and the predicted label corresponding to the feature vector of the high-value sample, the lightweight incremental learning module is incrementally trained at the edge end to update the module parameters of the lightweight incremental learning module; The updated module parameters at the edge are periodically synchronized with the main model in the cloud to complete the update of the business model.
7. A data management system based on multimodal loops, characterized in that, include: A data processing unit is used to acquire a multimodal business dataset from a business scenario; and to perform cross-modal feature alignment and fusion processing on the multimodal business data corresponding to each business task in the multimodal business dataset to obtain the feature vector to be evaluated corresponding to each business task; wherein, the multimodal business dataset includes multimodal business data corresponding to multiple business tasks; The multimodal feature pool includes a temporary storage area and a core storage area; the temporary storage area is used to store the feature vectors to be evaluated; the core storage area is used to store historical high-value feature vectors. The label prediction unit is used to generate predicted labels for each feature vector to be evaluated in the temporary storage area based on the historical high-value feature vectors in the core storage area, and to calculate the label confidence of each predicted label. The model update unit is used to filter out high-value sample feature vectors from the temporary storage area based on the label confidence level; and to update the business model used to perform business tasks using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The sample reflow unit is used to evaluate the value of the high-value sample feature vectors. If the value evaluation result meets the predetermined reflow conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
8. A service data management device based on multimodal loops, characterized in that, include: The data acquisition module is used to acquire multimodal business datasets from business scenarios; wherein, the multimodal business datasets include multimodal business data corresponding to multiple business tasks; The data processing module is used to perform cross-modal feature alignment and fusion processing on the multimodal business data corresponding to each business task in the multimodal business dataset, to obtain the feature vector to be evaluated corresponding to each business task, and to store each feature vector to be evaluated in the temporary storage area of the multimodal feature pool. The label prediction module is used to generate predicted labels for each feature vector to be evaluated in the temporary storage area based on historical high-value feature vectors stored in the core storage area of the multimodal feature pool, and to calculate the label confidence of the predicted labels. The model update module is used to filter out high-value sample feature vectors from the temporary storage area based on the label confidence level; and to update the business model used to perform business tasks using the high-value sample feature vectors and the predicted labels corresponding to the high-value sample feature vectors. The data backflow module is used to evaluate the value of the high-value sample feature vectors. If the value evaluation result meets the predetermined backflow conditions, the high-value sample feature vectors are transferred from the temporary storage area to the core storage area of the multimodal feature pool.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.