Man-machine collaborative decision-making method based on adaptive learning

By constructing a multimodal fusion layer and an adaptive learning layer, and combining reinforcement learning and meta-learning methods, the problem of doctor-model collaboration in pathological diagnosis models is solved, and efficient, reliable and flexible diagnostic results output for human-machine collaborative decision-making is achieved.

CN121902014APending Publication Date: 2026-04-21CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2025-12-05
Publication Date
2026-04-21

Smart Images

  • Figure CN121902014A_ABST
    Figure CN121902014A_ABST
Patent Text Reader

Abstract

The invention relates to the field of man-machine collaborative decision-making, and discloses a man-machine collaborative decision-making method based on adaptive learning, which comprises the steps of constructing a multi-modal fusion layer, a dynamic feedback regulation layer, an adaptive learning layer and a collaborative decision-making output layer, extracting pathological information features through the multi-modal fusion layer, and generating preliminary feature representation in combination with doctor operation feedback; a self-adaptive feedback signal and model confidence are introduced, and a dynamic regulation and control strategy is generated; performing real-time weight fusion on the doctor decision and the model output to form a weighted comprehensive decision, and dynamically updating the fusion weight according to the doctor confidence and the model performance; a parameter updating signal and strategy adjusting information are generated by adopting an ARCL algorithm integrating reinforcement learning and meta learning; and completing model parameter updating through the dynamic regulation and control module, and generating final decision output on the basis of the updated model parameters and the fusion weight. The method has the advantage that the human-machine cognitive alignment capability and the complex pathological diagnosis stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-machine collaborative decision-making, specifically a human-machine collaborative decision-making method based on adaptive learning. Background Technology

[0002] In the context of AI-enabled medical diagnosis, large-scale pathological models are gradually becoming an important support for intelligent diagnosis. However, existing models are mostly static learning-based, making it difficult to achieve effective collaboration between doctors and models in complex and dynamic clinical scenarios. Although traditional AI models can automatically identify pathological features, they lack a real-time response mechanism to doctors' decision-making behavior, resulting in an "information silo" phenomenon in human-computer interaction: on the one hand, the model output is difficult to explain the doctor's thinking logic and cannot be integrated into the clinical decision chain; on the other hand, the doctor's feedback information fails to effectively feed back into the model's learning, resulting in insufficient interpretability, stability, and trustworthiness of diagnostic results. Especially in renal pathology tasks, due to high case heterogeneity and blurred feature boundaries, static models are prone to misjudgment and decreased generalization ability in complex pathological identification. Existing systems generally lack adaptive learning and dynamic adjustment mechanisms, making it impossible to achieve real-time cognitive alignment and strategy co-evolution between "human-machine-task". Therefore, it is necessary to design an adaptive learning-based human-machine collaborative decision-making method to improve human-machine cognitive alignment and the stability of complex pathological diagnosis. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a human-machine collaborative decision-making method based on adaptive learning, which has the advantages of improving human-machine cognitive alignment and stability in complex pathological diagnosis, and solves the problems mentioned in the background technology.

[0004] To achieve the aforementioned goals of improving human-machine cognitive alignment and the stability of complex pathological diagnosis, this invention provides the following technical solution: a human-machine collaborative decision-making method based on adaptive learning, comprising the following steps: A multimodal fusion layer, a dynamic feedback control layer, an adaptive learning layer, and a collaborative decision-making output layer are constructed. The pathological information features are extracted through the multimodal fusion layer and combined with the doctor's operation feedback to generate preliminary features. Based on preliminary features, adaptive feedback signals and model confidence are introduced, and the learning rate is dynamically adjusted according to the task complexity to generate a dynamic control strategy. According to the dynamic adjustment strategy, the doctor's decision and the model output are fused in real time to form a weighted comprehensive decision, and the fusion weights are dynamically updated based on the doctor's confidence and the model performance. By combining the dynamically updated fusion weights, the ARCL algorithm, which integrates reinforcement learning and meta-learning, is adopted. Multi-objective optimization is performed by combining prediction loss, human-machine consistency loss and uncertainty regularization term to generate parameter update signals and policy adjustment information. Based on parameter update signals and strategy adjustment information, combined with doctor correction signals, confidence feedback and feature attention areas, the model parameters are updated through a dynamic adjustment module, and the final decision output is generated based on the updated model parameters and fusion weights.

[0005] Preferably, the process of constructing the multimodal fusion layer, dynamic feedback control layer, adaptive learning layer, and collaborative decision-making output layer is as follows: The pathological images, structured medical record information, and operation sequences are input into the multimodal coding module to extract a unified feature vector; A feedback control path is established based on the task flow, the adaptive learning layer is connected to the dynamic update interface, and a callable collaborative decision output unit is constructed.

[0006] Preferably, the process of generating preliminary features is as follows: The unified feature vector is fed into the encoding, alignment and feature extraction module to obtain the local and structural features of the pathological image; Synchronous feedback of doctor's operation signals in the control path transforms the area of ​​focus, annotation and correction behavior into auxiliary features; The input specifications and fusion characteristics of the collaborative decision-making unit are matched and spliced ​​to form preliminary features.

[0007] Preferably, the process of introducing adaptive feedback signals and model confidence is as follows: Based on preliminary features, the prediction results of the model are analyzed and confidence indices are calculated, including output distribution stability and class discrimination. Simultaneously, the doctor's operational feedback is screened and coded, transforming the doctor's focus, corrective actions, and changes in judgment into feedback signals describing operational tendencies; The model confidence vector and feedback signal are input together into the dynamic feedback control layer, and joint feedback is established through feature alignment and temporal correlation.

[0008] Preferably, the process of generating a dynamic control strategy is as follows: Based on joint feedback, the confidence level change trend, feedback signal strength, and characteristic distribution pattern are analyzed; Extract feature density, regional dissimilarity, boundary ambiguity, and decision offset to construct a task complexity index; Based on this indicator, the learning rate, parameter update step size, and internal control factor are adjusted synchronously to form a dynamic control strategy.

[0009] Preferably, the process of fusing doctor's decisions and model outputs through real-time weighting to form a weighted comprehensive decision is as follows: Based on the weight adjustment parameters provided by the dynamic control strategy, the structured decision signals given by doctors and the current prediction results of the model are quantified, normalized and weighted item by item. During the fusion process, synchronous weighted calculations are performed on different feature channels, time series nodes, and spatial areas of interest to generate a weighted comprehensive decision that can represent human-machine joint judgment.

[0010] Preferably, the process of dynamically updating the fusion weights based on doctor confidence and model performance is as follows: Based on weighted comprehensive decision-making, doctors' confidence scores for the current conclusion, historical performance indicators of the model for similar tasks, and stability measures of the current prediction are obtained in real time. These inputs are mapped to weight update rules in the dynamic control strategy; The human-machine fusion coefficient is adjusted in a fine-grained manner according to the rules, including the dynamic redistribution of cross-channel weights, time step weights, and decision region weights; After the weight update is completed, it will be synchronized with the current interaction cycle and written into the collaborative decision-making path.

[0011] Preferably, the ARCL algorithm, which integrates reinforcement learning and meta-learning, is as follows: After writing the updated fusion weight into the collaborative decision path, the fusion weight and the corresponding comprehensive decision are retrieved from the collaborative decision path. The current state features are input into the ARCL algorithm to construct the state vector, action space, and policy parameter set; In the reinforcement learning module, actions are generated based on the current policy, and interactive feedback is calculated. The meta-learning module simultaneously analyzes cross-task differences, adjusts the update step size of inner and outer loops, and constructs a transferable parameter structure. Through policy evaluation and gradient calculation, the policy gradient and parameter correction direction are obtained.

[0012] Preferably, the process of generating parameter update signals and strategy adjustment information is as follows: Based on the current policy gradient and parameter correction direction, the prediction bias, human-machine consistency difference and decision uncertainty measure are calculated, and the three types of measure information are mapped to a unified multi-task error signal space. The mapped multi-task error signals are input into the multi-objective optimization module. Within the module, the error signals are weighted, fused, and conflict-reconciled based on the model parameter structure, loss coupling relationship, and gradient sensitivity. Based on the correction results of the policy gradient by the multi-objective optimization module, the optimal adjustment range of the loss term and the internal weight allocation scheme are determined by solving the joint optimization function, and parameter update signals and corresponding human-machine collaborative policy adjustment signals are generated.

[0013] Preferably, the process of generating the final decision output based on the updated model parameters and fusion weights is as follows: The parameter update signal and the corresponding human-machine collaboration strategy adjustment signal are uniformly input into the dynamic control module; The model's parameters, attention distribution, feature weights, and human-machine fusion weights are updated layer by layer, while the internal control variables and channel activation paths are adjusted synchronously. The collaborative decision-making output structure calls the latest model weights and fusion coefficients to generate the final decision output for the current iteration cycle.

[0014] Compared with existing technologies, this invention provides a human-computer collaborative decision-making method based on adaptive learning, which has the following beneficial effects: This invention integrates pathological images, structured medical records, and physician feedback through multimodal information fusion, achieving a comprehensive representation of medical information and improving the accuracy and completeness of features. It dynamically regulates the learning process using adaptive feedback signals and model confidence, enabling the model to automatically adjust the learning rate and internal control parameters under different task complexities, thereby enhancing the flexibility and stability of decision-making. Real-time weight fusion and dynamic updates of fusion coefficients effectively combine physician experience with model predictions, improving the consistency and reliability of human-machine collaborative judgment. Combining reinforcement learning and meta-learning for strategy optimization, and considering prediction bias, human-machine consistency, and uncertainty through multi-objective optimization, it generates refined parameter update and strategy adjustment schemes, enhancing the model's adaptability to complex scenarios. Finally, a dynamic control module updates parameters and fusion weights, ensuring that the decision output after each iteration is accurate, traceable, and highly coordinated with physician operations, thus significantly improving decision-making efficiency, reliability, and adaptability to different task conditions. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the method of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1: Please refer to Figure 1 As shown in the embodiment of the present invention, a human-machine collaborative decision-making method based on adaptive learning includes the following steps: S1: Construct a multimodal fusion layer, a dynamic feedback control layer, an adaptive learning layer, and a collaborative decision-making output layer. The multimodal fusion layer extracts pathological information features and combines them with doctor's operational feedback to generate preliminary features.

[0018] The process of constructing the multimodal fusion layer, dynamic feedback control layer, adaptive learning layer, and collaborative decision-making output layer in S1 is as follows: The pathological images, structured medical record information, and operation sequences are input into the multimodal coding module to extract a unified feature vector; The pathological images are preprocessed, including resolution unification, grayscale standardization, and Gaussian filtering for noise reduction. Local texture, edge, and structural features are extracted using a convolutional neural network to generate a two-dimensional feature map. Structured medical record information, such as laboratory indicators, diagnostic conclusions, and clinical description text, is converted into a numerical vector compatible with image feature vectors through discretization encoding, normalization, and word embedding. Operation sequence data, including the doctor's annotation position, correction actions, and dwell time, are generated into a continuous vector through time series encoding. The multimodal encoding module extracts features from the three types of inputs respectively, and then forms a unified feature vector through vector concatenation and channel weighted fusion.

[0019] A feedback control path is established based on the task flow, the adaptive learning layer is connected to the dynamic update interface, and a callable collaborative decision output unit is constructed. Based on the operational process of the diagnostic task, a feedback control path is constructed, and the doctor's operation signals, task stage information and feature vectors are uniformly mapped to the dynamic feedback framework. The adaptive learning layer accesses the feedback framework through a preset dynamic update interface, obtains updated data in real time, and iteratively adjusts the model parameters according to the learning rate scheduling strategy and weight update rules, including synchronously updating the weighting coefficients of feature vectors, attention distribution and channel activation paths.

[0020] The process of generating preliminary features in S1 is as follows: The unified feature vector is fed into the encoding, alignment and feature extraction module to obtain the local and structural features of the pathological image; Local texture features of pathological images are extracted through convolution and pooling operations. At the same time, the overall structural information is obtained through a multi-scale feature extraction structure module. Then, feature standardization is performed through batch normalization and nonlinear activation functions to maintain the consistency of features of different modalities. Subsequently, each feature channel is aligned according to spatial location and hierarchical information to ensure that the local features of the image correspond to the overall structural features in the same space. Finally, a high-dimensional vector reflecting key morphology, boundary and texture patterns is extracted through a feature extraction module.

[0021] Synchronous feedback of doctor's operation signals in the control path transforms the area of ​​focus, annotation and correction behavior into auxiliary features; In the feedback control path, doctor operation signals are collected in real time, including annotation locations, correction behaviors, and areas of attention. These operation signals are mapped to the same spatial dimension as the model feature vector to construct auxiliary feature vectors, where attention weights or correction intensities for specific regions in each dimension are defined. During the mapping process, the signal amplitude is adjusted through linear or nonlinear projection functions, and time series encoding is combined to preserve the operation sequence information, enabling the model to synchronously acquire the doctor's attention to and correction actions on local features.

[0022] The input specifications and fusion features of the collaborative decision-making unit are matched and combined to form preliminary features; The image feature vectors extracted by encoding are aligned and concatenated with the doctor's operation auxiliary features according to the input specifications of the collaborative decision-making unit. Weighted matching is performed on each feature channel and spatial location to maintain the correspondence between image features and operation features. At the same time, amplitude differences are eliminated through normalization and channel standardization to finally generate preliminary features.

[0023] S2: Based on the initial features, an adaptive feedback signal and model confidence are introduced, and the learning rate is dynamically adjusted according to the task complexity to generate a dynamic control strategy.

[0024] The process of introducing adaptive feedback signals and model confidence in S2 is as follows: Based on preliminary features, the prediction results of the model are analyzed and confidence indices are calculated, including output distribution stability and class discrimination. Preliminary features are input into the prediction model, which outputs the probability distribution of each category and the activation value of each prediction result. By performing statistical analysis on the output probability distribution, the distribution stability and category discrimination index of each category are calculated. The distribution stability is quantified by the difference between historical prediction results and current prediction results. The category discrimination index is calculated by the difference between the maximum probability and the second maximum probability or the entropy value. During the calculation process, the output value of each channel is normalized, and the confidence index is vectorized to form the model confidence vector.

[0025] Simultaneously, the doctor's operational feedback is screened and coded, transforming the doctor's focus, corrective actions, and changes in judgment into feedback signals describing operational tendencies; Doctor operation data is obtained from the feedback control path, including operation timestamps, coordinates of the area of ​​interest, annotation type, and correction behavior. The raw operation data is cleaned to remove duplicate or abnormal actions. Then, each feedback is mapped to a numerical feature or vector representation according to the operation type. The area of ​​interest is mapped to the feature weight by its position in the feature space. The annotation and correction behavior are transformed into auxiliary vectors through a set of predefined encoding rules.

[0026] The model confidence vector and the feedback signal are jointly input into the dynamic feedback control layer, and joint feedback is established through feature alignment and temporal correlation. The encoded doctor operation feedback vector is combined with the model confidence vector. The information of each channel is aligned in a unified feature space through a feature alignment mechanism. Combined with the temporal correlation method, the feedback signals at different time points are matched with the corresponding predicted confidence to form a multi-dimensional joint feedback.

[0027] The process of generating dynamic control strategies in S2 is as follows: Based on joint feedback, the confidence level change trend, feedback signal strength, and characteristic distribution pattern are analyzed; The joint feedback vector is input into the analysis module, and time series analysis is performed on the model confidence vector and the doctor's operation feedback vector respectively. The confidence change trend is obtained by calculating the difference magnitude and stability index between continuous prediction results. The feedback signal strength is obtained by statistically quantifying the weight of the attention area and the frequency of correction actions. The feature distribution pattern is formed by density estimation and statistical analysis of the numerical distribution of different feature channels, forming the spatial and channel distribution characteristics of each feature.

[0028] Extract feature density, regional dissimilarity, boundary ambiguity, and decision offset to construct a task complexity index; Based on the analyzed features, the local feature density is calculated, which is the degree of feature value clustering in a specific spatial region, to assess whether the information is concentrated or sparse. The regional difference is quantified by comparing the mean and variance of the feature distribution in different regions to measure the difference between local and global information. The boundary ambiguity is measured by statistically analyzing the changes in edge response or feature gradient to measure structural clarity. The decision offset is calculated by the difference between historical predictions and current prediction results to quantify the uncertainty or bias of the model output. All indicators are mapped to a task complexity vector to form a quantitative standard.

[0029] Based on this indicator, the learning rate, parameter update step size, and internal control factor are adjusted synchronously to form a dynamic control strategy; The task complexity vector is input into the dynamic adjustment module, which adjusts the hyperparameters of the model training in real time based on the values ​​of each indicator. The learning rate is updated according to a rule that is inversely proportional to the task complexity. That is, the learning rate is reduced when the task complexity is high to ensure training stability, and increased when the complexity is low to accelerate convergence. The parameter update step size is achieved by adjusting the gradient step size or the weight update magnitude, which corresponds to the task complexity and feature change trend. The internal control factors include attention channel gain, feature channel weight, and activation path adjustment coefficient. By calculating the importance of each channel in the joint feedback and redistributing the weights, the feature update of the model is synchronized under dynamic tasks, generating a complete dynamic adjustment strategy.

[0030] S3: Based on the dynamic adjustment strategy, the doctor's decision and the model output are fused in real time to form a weighted comprehensive decision, and the fusion weights are dynamically updated according to the doctor's confidence and model performance.

[0031] In S3, the process of fusing physician decisions and model outputs in real-time weighted analysis to form a weighted comprehensive decision is as follows: Based on the weight adjustment parameters provided by the dynamic control strategy, the structured decision signals given by doctors and the current prediction results of the model are quantified, normalized and weighted item by item. The weight parameters generated by the dynamic adjustment strategy are input into the fusion module. The doctor's decision signals include diagnosis selection, labeled region and correction operation sequence. The model prediction results include various prediction probabilities and feature channel responses. The two types of inputs are numerically quantified at the channel and time nodes, the qualitative operations are transformed into computable vectors, and the normalization method is used to unify them to the same numerical range. Then, the normalized vectors are weighted according to the dynamic weight parameters.

[0032] During the fusion process, synchronous weighted calculations are performed on different feature channels, time series nodes, and spatial areas of interest to generate a weighted comprehensive decision that can represent human-machine joint judgment. During the fusion process, weighting operations are performed on each feature channel, each time series node, and each spatial region of interest. Feature channel weighting is used to balance the contribution of different modalities or different types of features to the final decision. Time series node weighting ensures the synchronicity and historical consistency between the model and the doctor's decision in continuous interaction. Spatial region of interest weighting compares the doctor's annotations and the model's feature response regions and assigns corresponding weights to achieve fine fusion of regional information. Each weighting operation is completed through matrix multiplication or tensor operations, and the channel, time, and spatial indices are preserved during the calculation process. After synchronous weighting calculation, all weighted vectors are summarized and integrated in the channel, time, and spatial dimensions to form the final weighted comprehensive decision vector.

[0033] The process of dynamically updating the fusion weights based on doctor confidence and model performance in S3 is as follows: Based on weighted comprehensive decision-making, doctors' confidence scores for the current conclusion, historical performance indicators of the model for similar tasks, and stability measures of the current prediction are obtained in real time. After completing the weighted comprehensive decision, the doctor's operational behavior in the current diagnostic task is collected in real time, including the selection preference of each diagnostic option, the degree of confirmation of the marked area, and the correction actions of the previous prediction results. At the same time, the timing, duration and spatial score of the doctor's operation are recorded. These operational data are encoded and standardized to form a computable confidence vector.

[0034] These inputs are mapped to weight update rules in the dynamic control strategy; Simultaneously, the model's performance metrics in similar historical tasks are acquired, including prediction accuracy, output category distribution stability, feature channel response strength, and fluctuations in continuous prediction. The current prediction probability distribution and channel activation mode are recorded synchronously to form a model performance vector. These vector data, together with the doctor's confidence vector, are used as inputs to map to the dynamic adjustment strategy module. According to the preset weight update rules, the input variables are transformed into adjustment schemes for human-machine fusion coefficients, including calculation formulas and adjustment values ​​for cross-feature channel weights, time node weights, and key decision region weights.

[0035] The human-machine fusion coefficient is adjusted in a fine-grained manner according to the rules, including the dynamic redistribution of cross-channel weights, time step weights, and decision region weights; According to the calculated update scheme, the human-machine fusion coefficient is adjusted in a fine-grained manner layer by layer, channel by channel, and time node by time node. The contributions of different channels and time steps are redistributed, and the weights of each key decision area are dynamically redistributed, so that the integrated decision after fusion can reflect the consistency between the doctor's operational intention and the reliability of the model prediction in different dimensions.

[0036] After the weight update is completed, it will be synchronized with the current interaction cycle and written into the collaborative decision-making path; After the update is completed, the new fusion weights are aligned with the input data of the current interaction cycle in terms of time and task stage, and written into the collaborative decision path.

[0037] S4: Combining the dynamically updated fusion weights, the ARCL algorithm, which integrates reinforcement learning and meta-learning, is used to perform multi-objective optimization by combining prediction loss, human-machine consistency loss, and uncertainty regularization term, thereby generating parameter update signals and policy adjustment information.

[0038] The process of using the ARCL algorithm, which combines reinforcement learning and meta-learning, in S4 is as follows: After writing the updated fusion weight into the collaborative decision path, the fusion weight and the corresponding comprehensive decision are retrieved from the collaborative decision path. The fusion weights calculated for each subtask or submodel in the previous training or decision round are organized into a fusion weight table according to the time series and task identifier, and written into the buffer of the collaborative decision path. The buffer is a multi-dimensional array structure, and the index includes the task number, timestamp, and status identifier. After writing, the corresponding fusion weight record is retrieved from the collaborative decision path according to the identifier and status of the current task to be processed, and matched with the previously saved comprehensive decision vector to form a set of fusion weight-decision pairs input to the ARCL algorithm.

[0039] The current state features are input into the ARCL algorithm to construct the state vector, action space, and policy parameter set; The current state features, including continuous and discrete variables, are obtained from the environmental perception module and the task monitoring module. The state features are vectorized and encoded and arranged in a predefined order to form a complete state vector. Each element in the state vector corresponds to a specific environmental parameter or task indicator. The action space is generated according to the current task category and constraints, including possible decision actions and their parameter ranges. The action space is represented in a mixed form of discrete and continuous variables and stored in the action dictionary. The policy parameter set consists of the weight matrix, bias vector, and possible gating parameters of the current policy network. The state vector, action space, and policy parameter set are provided as input to the execution engine of the ARCL algorithm.

[0040] In the reinforcement learning module, actions are generated based on the current policy, and interactive feedback is calculated. The policy network performs forward propagation on the input state vector to generate action distribution or action value vector, and selects specific actions through action selection function, including random sampling or greedy selection methods. The selected actions are applied to environmental simulation or actual tasks, and the environmental response and task feedback after the action is executed are recorded, including reward signals, state transition information and observation results. Then, the execution results are calculated with a predefined reward function to obtain the immediate reward value and cumulative reward.

[0041] The meta-learning module simultaneously analyzes cross-task differences, adjusts the update step size of inner and outer loops, and constructs a transferable parameter structure. Feature differences are extracted from historical task data and current task execution results, including state feature distribution shifts, action selection probability changes, and reward curve differences. Gradient information for each task is statistically analyzed, and gradient correlation and sensitivity indices between tasks are calculated. The learning rate and update step size of the meta-learning inner and outer loops are adjusted based on the analysis results, and a transferable parameter structure is constructed.

[0042] After policy evaluation and gradient calculation, the policy gradient and parameter correction direction are obtained; Using the current policy, forward inference is performed on the state vector to generate actions and calculate corresponding rewards. Then, in the reinforcement learning module, the gradient of each policy parameter is calculated using the policy gradient algorithm, including the partial derivative of the action probability distribution with respect to the reward. At the same time, the inter-task gradient correction information provided by the meta-learning module is combined to weight or correct the policy gradient to form the final parameter update direction.

[0043] The process of generating parameter update signals and policy adjustment information in S4 is as follows: Based on the current policy gradient and parameter correction direction, the prediction bias, human-machine consistency difference and decision uncertainty measure are calculated, and the three types of measure information are mapped to a unified multi-task error signal space. The algorithm reads the policy gradient vector output by the reinforcement learning module and the parameter correction direction vector provided by the meta-learning module, calculates the deviation in predicted reward for each state-action pair, and forms a prediction deviation matrix. At the same time, it obtains the operation behavior sequence from the historical human-machine collaboration record, compares the current action selection with the human operation trajectory, and calculates the human-machine consistency difference matrix. Then, it calculates the decision uncertainty measurement matrix based on the action probability distribution and the uncertainty index of state transition. After numerical standardization and dimensional unification, the three types of matrices are mapped to the same-dimensional multi-task error signal space using a preset mapping function, forming an error signal vector that can be directly input into the optimization module.

[0044] The mapped multi-task error signals are input into the multi-objective optimization module. Within the module, the error signals are weighted, fused, and conflict-reconciled based on the model parameter structure, loss coupling relationship, and gradient sensitivity. A uniform dimension error signal vector is input into the multi-objective optimization module. Within the module, a weight coefficient matrix is ​​established for each error signal. This matrix is ​​calculated based on task priority, gradient sensitivity, and historical error stability. Subsequently, a weighted summation operation is performed on the error signals. When conflicting error signals are encountered, the least squares harmonic method is used to calculate the weighted correction, so that the error gradient direction remains as consistent as possible across multiple tasks.

[0045] Based on the correction results of the policy gradient by the multi-objective optimization module, the optimal adjustment range of the loss term and the internal weight allocation scheme are determined by solving the joint optimization function, and parameter update signals and corresponding human-machine collaborative policy adjustment signals are generated. The correction policy gradient vector output by the multi-objective optimization module is read and combined with the current policy gradient and gradient constraints to construct a joint optimization function. The joint optimization function includes a weighted summation term of the loss functions of each task, a gradient regularization term, and a parameter constraint term. The joint optimization function is optimized using an iterative solution method. In each iteration, the internal weight allocation scheme is dynamically updated according to the gradient direction and magnitude, and the optimal adjustment magnitude of each loss term is calculated. Finally, after the solution is completed, a parameter update signal matrix is ​​generated.

[0046] S5: Based on parameter update signals and strategy adjustment information, combined with doctor correction signals, confidence feedback and feature attention areas, the model parameters are updated through the dynamic adjustment module, and the final decision output is generated based on the updated model parameters and fusion weights.

[0047] In S5, the process of generating the final decision output based on the updated model parameters and fusion weights is as follows: The parameter update signal and the corresponding human-machine collaboration strategy adjustment signal are uniformly input into the dynamic control module; The parameter update signal matrix and human-machine collaboration strategy adjustment information generated in the previous stage are sorted and uniformly encoded according to the task sequence and timestamp to form a standardized input tensor. The tensor contains the update magnitude, direction and corresponding task state index of each strategy parameter, and records the fine-tuning value and weight allocation information of the human-machine collaboration adjustment. After receiving the tensor, the dynamic control module parses each record and generates the parameter update instruction queue inside the module.

[0048] The model's parameters, attention distribution, feature weights, and human-machine fusion weights are updated layer by layer, while the internal control variables and channel activation paths are adjusted synchronously. The dynamic adjustment module traverses each network layer of the model from high to low according to the input update instruction queue, and performs incremental updates on the convolutional kernel weights, fully connected layer weights, attention mechanism weights, and feature channel activation intensities. During the update process, the module refers to the adjustment values ​​of the human-machine collaboration strategy to fine-tune the activation path of each channel, so that the output amplitude and weight distribution of specific feature channels are consistent with the update signal. The module also updates internal control variables, including gating coefficients, batch normalization parameters, and feature response thresholds, to ensure that parameter updates are reflected synchronously in forward computation and gradient propagation, and to maintain the consistency of internal model computation.

[0049] The collaborative decision-making output structure calls the latest model weights and fusion coefficients to generate the final decision output for the current iteration cycle; After completing the layer-by-layer update, the collaborative decision output structure integrates the parameters at each level of the model, the attention distribution, and the human-machine fusion weights. It then performs forward propagation calculations according to the predefined calculation path. The collaborative decision output structure selects the corresponding action candidate set based on the task identifier and performs weighted summation or weighted selection on each candidate action in combination with the latest fusion weights to generate the final decision vector.

[0050] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A human-computer collaborative decision-making method based on adaptive learning, characterized in that, Includes the following steps: A multimodal fusion layer, a dynamic feedback control layer, an adaptive learning layer, and a collaborative decision-making output layer are constructed. The pathological information features are extracted through the multimodal fusion layer and combined with the doctor's operation feedback to generate preliminary features. Based on preliminary features, adaptive feedback signals and model confidence are introduced, and the learning rate is dynamically adjusted according to the task complexity to generate a dynamic control strategy. According to the dynamic adjustment strategy, the doctor's decision and the model output are fused in real time to form a weighted comprehensive decision, and the fusion weights are dynamically updated based on the doctor's confidence and the model performance. By combining the dynamically updated fusion weights, the ARCL algorithm, which integrates reinforcement learning and meta-learning, is adopted. Multi-objective optimization is performed by combining prediction loss, human-machine consistency loss and uncertainty regularization term to generate parameter update signals and policy adjustment information. Based on parameter update signals and strategy adjustment information, combined with doctor correction signals, confidence feedback and feature attention areas, the model parameters are updated through a dynamic adjustment module, and the final decision output is generated based on the updated model parameters and fusion weights.

2. The human-machine collaborative decision-making method based on adaptive learning according to claim 1, characterized in that, The process of constructing the multimodal fusion layer, dynamic feedback control layer, adaptive learning layer, and collaborative decision-making output layer is as follows: The pathological images, structured medical record information, and operation sequences are input into the multimodal coding module to extract a unified feature vector; A feedback control path is established based on the task flow, the adaptive learning layer is connected to the dynamic update interface, and a callable collaborative decision output unit is constructed.

3. The human-computer collaborative decision-making method based on adaptive learning according to claim 2, characterized in that, The process of generating preliminary features is as follows: The unified feature vector is fed into the encoding, alignment and feature extraction module to obtain the local and structural features of the pathological image; Synchronous feedback of doctor's operation signals in the control path transforms the area of ​​focus, annotation and correction behavior into auxiliary features; The input specifications and fusion characteristics of the collaborative decision-making unit are matched and spliced ​​to form preliminary features.

4. The human-computer collaborative decision-making method based on adaptive learning according to claim 3, characterized in that, The process of introducing adaptive feedback signals and model confidence is as follows: Based on preliminary features, the prediction results of the model are analyzed and confidence indices are calculated, including output distribution stability and class discrimination. Simultaneously, the doctor's operational feedback is screened and coded, transforming the doctor's focus, corrective actions, and changes in judgment into feedback signals describing operational tendencies; The model confidence vector and feedback signal are input together into the dynamic feedback control layer, and joint feedback is established through feature alignment and temporal correlation.

5. The human-machine collaborative decision-making method based on adaptive learning according to claim 4, characterized in that, The process of generating a dynamic control strategy is as follows: Based on joint feedback, the confidence level change trend, feedback signal strength, and characteristic distribution pattern are analyzed; Extract feature density, regional dissimilarity, boundary ambiguity, and decision offset to construct a task complexity index; Based on this indicator, the learning rate, parameter update step size, and internal control factor are adjusted synchronously to form a dynamic control strategy.

6. The human-computer collaborative decision-making method based on adaptive learning according to claim 5, characterized in that, The process of fusing physician decisions and model outputs in real time to form a weighted comprehensive decision is as follows: Based on the weight adjustment parameters provided by the dynamic control strategy, the structured decision signals given by doctors and the current prediction results of the model are quantified, normalized and weighted item by item. During the fusion process, synchronous weighted calculations are performed on different feature channels, time series nodes, and spatial areas of interest to generate a weighted comprehensive decision that can represent human-machine joint judgment.

7. The human-machine collaborative decision-making method based on adaptive learning according to claim 6, characterized in that, The process of dynamically updating the fusion weights based on doctor confidence and model performance is as follows: Based on weighted comprehensive decision-making, doctors' confidence scores for the current conclusion, historical performance indicators of the model for similar tasks, and stability measures of the current prediction are obtained in real time. These inputs are mapped to weight update rules in the dynamic control strategy; The human-machine fusion coefficient is adjusted in a fine-grained manner according to the rules, including the dynamic redistribution of cross-channel weights, time step weights, and decision region weights; After the weight update is completed, it will be synchronized with the current interaction cycle and written into the collaborative decision-making path.

8. The human-computer collaborative decision-making method based on adaptive learning according to claim 7, characterized in that, The process of using the ARCL algorithm, which integrates reinforcement learning and meta-learning, is as follows: After writing the updated fusion weight into the collaborative decision path, the fusion weight and the corresponding comprehensive decision are retrieved from the collaborative decision path. The current state features are input into the ARCL algorithm to construct the state vector, action space, and policy parameter set; In the reinforcement learning module, actions are generated based on the current policy, and interactive feedback is calculated. The meta-learning module simultaneously analyzes cross-task differences, adjusts the update step size of inner and outer loops, and constructs a transferable parameter structure. Through policy evaluation and gradient calculation, the policy gradient and parameter correction direction are obtained.

9. A human-computer collaborative decision-making method based on adaptive learning according to claim 8, characterized in that, The process of generating parameter update signals and policy adjustment information is as follows: Based on the current policy gradient and parameter correction direction, the prediction bias, human-machine consistency difference and decision uncertainty measure are calculated, and the three types of measure information are mapped to a unified multi-task error signal space. The mapped multi-task error signals are input into the multi-objective optimization module. Within the module, the error signals are weighted, fused, and conflict-reconciled based on the model parameter structure, loss coupling relationship, and gradient sensitivity. Based on the correction results of the policy gradient by the multi-objective optimization module, the optimal adjustment range of the loss term and the internal weight allocation scheme are determined by solving the joint optimization function, and parameter update signals and corresponding human-machine collaborative policy adjustment signals are generated.

10. A human-computer collaborative decision-making method based on adaptive learning according to claim 9, characterized in that, The process of generating the final decision output based on the updated model parameters and fusion weights is as follows: The parameter update signal and the corresponding human-machine collaboration strategy adjustment signal are uniformly input into the dynamic control module; The model's parameters, attention distribution, feature weights, and human-machine fusion weights are updated layer by layer, while the internal control variables and channel activation paths are adjusted synchronously. The collaborative decision-making output structure calls the latest model weights and fusion coefficients to generate the final decision output for the current iteration cycle.