Large model driven evolution of organic sensor array high fidelity information decoupling method

CN122734445APending Publication Date: 2026-09-11UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610888393.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

在面临复杂的生化测试环境、基线漂移或样本分布偏移时,模型极易因缺乏底层物理约束而出现过拟合,泛化能力衰退,难以满足高可靠性连续健康监测系统对算法透明度与安全性的要求

Benefits of technology

(1)突破了传统黑盒模型的不可解释性,实现了具有显式物理意义的高保真特征提取。本发明创新性地引入大语言模型作为语义感知的遗传算子,在一个开放式的符号代码空间中引导进化搜索,能够自主发现并提取反映有机电化学晶体管潜在电化学动力学的高阶物理特征,摆脱了繁重的人工物理校准,确保了特征提取过程的完全透明性和物理一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734445A_ABST
    Figure CN122734445A_ABST
Patent Text Reader

Abstract

This application relates to a high-fidelity information decoupling method for organic sensor arrays driven by a large model. The method includes: acquiring multi-channel response sequences of the organic sensor array and constructing a high-dimensional signal matrix; constructing a feature generation module driven by a large language model, using the large language model as a semantic-aware genetic operator, and generating physically interpretable feature extraction scripts through logical-aware mutation and structural-functional cross-evolution; constructing a meta-learner optimization module driven by the large language model, evolving to generate a meta-learner for adaptively allocating multi-channel decision weights; deploying candidate scripts in an isolated sandbox, combining a multi-objective hybrid evaluation function for security verification and quantification scoring, and driving model convergence through closed-loop feedback; and outputting high-fidelity decoupling results using an optimal signal decoupling model. This achieves high-precision, highly robust, and physically interpretable signal decoupling of organic sensor arrays.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the interdisciplinary fields of flexible bioelectronics, intelligent sensing and artificial intelligence, and in particular to a high-fidelity information decoupling method for organic sensing arrays driven by a large model. Background Technology

[0002] Vertical organic electrochemical transistors (vOECTs), with their volumetric gate control mechanism and ion-electron hybrid conduction characteristics, exhibit high transconductance and excellent signal amplification performance at extremely low operating voltages, making them an ideal hardware platform for constructing high-density, multi-channel biosensor arrays. Currently, in the field of vOECT multidimensional signal decoupling and biochemical sensing data processing, technical solutions have mainly evolved through the following stages: First, the physical mechanism modeling stage. Early approaches mainly involved constructing simplified electrochemical theoretical models to systematically analyze ion-sensing mechanisms and transfer characteristic curves, using this as the basis for quantitative signal monitoring and decoupling.

[0003] Second, the structural optimization and steady-state extremum extraction stage. As devices move towards vertical structures, their channel lengths shorten to the sub-micron level, resulting in faster transient response speeds. Existing conventional processing methods mostly rely on extracting the extrema of the transfer characteristic curve at a specific scan voltage or the static slope at a specific point as decoupling features.

[0004] Third, the pure data-driven algorithm stage. To address the nonlinearity of multi-channel signals in biochemical sensing, researchers have introduced various machine learning and neural network methods. For example, they have used multilayer perceptrons to fit nonlinear data, or used pooled computation combined with random forest algorithms to perform pattern recognition and classification on dynamic time-series sensing information.

[0005] Although the above methods have achieved some progress in specific experimental environments, the following unavoidable drawbacks still exist in the large-scale application of multi-channel vOECT arrays in practice: (1) Lack of physical mechanism constraints leads to low generalization ability and reliability of "black box" models. Existing pure data-driven models are essentially end-to-end "black box" structures, and their internal parameters do not have explicit physical and electrochemical meaning. When faced with complex biochemical testing environments, baseline drift, or sample distribution shifts, the models are prone to overfitting due to the lack of underlying physical constraints, resulting in a decline in generalization ability and making it difficult to meet the requirements of high-reliability continuous health monitoring systems for algorithm transparency and security.

[0006] (2) It is difficult to overcome the inherent heterogeneity between devices, which makes static calibration very prone to failure. Due to the randomness of the fabrication process of vertical devices, the devices in each channel of the array exhibit severe nonlinear heterogeneity. Traditional global mechanism models and mapping methods based on single static features are easily affected by the heterogeneity between devices because they cannot adaptively correct the unique deviation of each channel. This causes the overall calibration algorithm to fail when deployed across devices or on a large scale.

[0007] (3) High-order transient dynamic characteristics suffer severe loss and lack cross-channel collaborative decoupling mechanisms. The ion-electric coupling process within vOECT is accompanied by highly nonlinear transient dynamic responses. Existing dimensionality reduction schemes often extract only static scalars or isolated peaks, truncating the continuous physical dynamic information in the time-series response. At the same time, existing schemes often treat each sensor channel as an isolated data source and process them separately, failing to build an array-level collaborative sensing mechanism. This results in the system being unable to adaptively identify and suppress abnormal channel data caused by local hardware damage or environmental noise, limiting the upper limit of the overall array decoupling accuracy.

[0008] Therefore, there is an urgent need in related technologies for a way to neutralize the heterogeneity deviation between devices caused by the manufacturing process of organic sensor arrays, and to fully extract the transient response dynamic characteristics, so as to achieve high-fidelity decoupling of complex multi-channel electrical response signals, while taking into account both the physical interpretability of the algorithm and computational security. Summary of the Invention

[0009] Therefore, it is necessary to provide a high-fidelity information decoupling method for organic sensing arrays driven by large models to address the aforementioned technical problems.

[0010] Firstly, this application provides a high-fidelity information decoupling method for organic sensing arrays driven by large-model evolution. The method includes: Obtain the multi-channel sensor response sequence of an organic sensor array and construct a high-dimensional signal matrix; A feature generation module driven by a large language model is constructed. The large language model is used as a semantic-aware genetic operator. Logical-aware mutation and structural-functional crossover are performed in the symbolic code space to evolve and generate physically interpretable feature extraction scripts for each independent sensing channel, and the high-dimensional signal matrix is ​​mapped to low-dimensional physical features. A large language model-driven meta-learner optimization module is constructed, which uses the large language model to evolve and generate a meta-learner. The meta-learner is used to adaptively allocate multi-channel decision weights. The feature extraction script and the meta-learner generated by evolution are deployed in an isolated computing sandbox environment. Security verification and quantitative scoring are performed by combining a multi-objective hybrid evaluation function. The feature extraction script and the meta-learner are driven to converge toward the optimal signal decoupling model through closed-loop feedback. The optimal signal decoupling model, after training, is used to decouple the multi-channel response sequence of the organic sensor array, and outputs a high-fidelity decoupling result of the target marker concentration.

[0011] Optionally, in one embodiment of this application, the organic sensing array is a vertical organic electrochemical transistor array.

[0012] Optionally, in one embodiment of this application, the feature generation module driven by the large language model performs population initialization through a domain-specific prompt word template. The prompt word template encapsulates the electrochemical physics prior knowledge of the vertical organic electrochemical transistor into structured instructions, guiding the large language model to generate an initial feature extraction script in a physically constrained symbolic space.

[0013] Optionally, in one embodiment of this application, the logically-aware variation includes: The code and performance evaluation score of the selected well-performing and runnable parent feature extraction script are input into the large language model as semantic feedback context. The large language model then extends the parent feature extraction logic in mathematical or physical dimensions to generate a logically optimized offspring mutation script.

[0014] Optionally, in one embodiment of this application, the structural functional crossover includes: The code of the two parent feature extraction scripts is simultaneously input into the large language model. The large language model analyzes and reorganizes the feature extraction logic of the two parent scripts at the semantic level, retains complementary computational operators, removes redundant features, resolves naming conflicts and function dependencies, and generates a child cross script that integrates advantageous logic.

[0015] Optionally, in one embodiment of this application, the large language model-driven meta-learner optimization module further includes: The low-dimensional physical features are input into the base learner to generate concentration decoupling values ​​for each independent sensing channel, and the concentration decoupling values ​​of each channel are concatenated to construct a decoupling matrix. The large language model is used to perform meta-learner evolution in a preset regression algorithm cluster space. The evolution includes the evolution of algorithm type and the evolution of algorithm hyperparameters. The meta-learner takes the decoupling matrix as input and outputs the global concentration decoupling value after allocating decision weights for each channel.

[0016] Optionally, in one embodiment of this application, the base learners of each independent sensing channel are trained independently and do not share parameters. The algorithm type and hyperparameter configuration of each base learner are autonomously evolved by the large language model for the electrical response characteristics of the corresponding channel.

[0017] Optionally, in one embodiment of this application, the isolated computing sandbox is a multi-process isolated sandbox with computing resource limits and timeout interruption mechanisms, used to dynamically compile the feature extraction script and the meta-learner, intercept illegal calls and resource overruns, and ensure the safe operation of the evolution engine.

[0018] Optionally, in one embodiment of this application, the multi-objective hybrid evaluation function comprehensively measures the mean absolute error, mean absolute percentage error, coefficient of determination, and cross-channel stability index, and outputs a comprehensive evaluation score as the sole metric guiding the closed-loop feedback.

[0019] Optionally, in one embodiment of this application, the method further includes: A non-negative truncation process is introduced at the model output to force the negative values ​​in the original values ​​of the regression model output to be corrected to zero, ensuring that the output concentration meets the physical benchmark of electrochemical detection.

[0020] The above-mentioned high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution has the following advantages compared with existing technologies: (1) It breaks through the uninterpretability of traditional black-box models and realizes high-fidelity feature extraction with explicit physical meaning. This invention innovatively introduces a large language model as a semantic-aware genetic operator, which guides evolutionary search in an open symbolic code space. It can autonomously discover and extract high-order physical features that reflect the potential electrochemical dynamics of organic electrochemical transistors, get rid of the heavy manual physical calibration, and ensure the complete transparency and physical consistency of the feature extraction process.

[0021] (2) It effectively overcomes the inherent heterogeneity between array devices and significantly improves the decoding robustness of cross-channel signals. This invention constructs a topology-aware two-stage Stacking integrated learning architecture. In the first stage, it autonomously extracts highly decoupled features for independent channels to resist local baseline drift. In the second stage, it dynamically evolves the topology of the meta-learner to aggregate multi-channel signals, suppressing local hardware deviations and spatial distribution differences of sensors caused by inconsistent manufacturing processes. This ensures that the marker concentration can still be output stably and accurately in complex physiological or industrial environments.

[0022] (3) It balances the execution security of the algorithm in edge deployment with extremely high decoupling accuracy. This invention designs an isolated sandbox evaluation and a multi-dimensional hybrid fitness feedback mechanism to conduct rigorous execution security and resource constraint tests on all generated candidate code features, effectively avoiding system crashes caused by abnormal code. In the validation of array datasets with multiple different ions, the framework successfully reduced the decoupling error to an extremely low level, with the coefficient of determination reaching or approaching 1, providing a solid, secure and scalable algorithmic foundation for next-generation intelligent biochemical sensing and high-reliability health informatics. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating a high-fidelity information decoupling method for organic sensor arrays driven by a large model in one embodiment. Figure 2 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0025] This embodiment uses a vertical organic electrochemical transistor (vOECT) array as a specific implementation of an organic sensing array to describe in detail the specific implementation process of the method of the present invention.

[0026] In one embodiment, such as Figure 1 As shown, a high-fidelity information decoupling method for organic sensing arrays driven by large-scale model evolution is provided, including the following steps: S101: Obtain the multi-channel sensor response sequence of the organic sensor array and construct a high-dimensional signal matrix.

[0027] In this embodiment, firstly, the multi-channel sensor response sequence of the vOECT array is acquired, preprocessed, and a high-dimensional signal matrix is ​​constructed. Specifically, under a given biochemical sensing test environment, a predetermined dynamic gate voltage scan sequence is applied to the vOECT array. The range of the dynamic gate voltage scan sequence is set to -0.4 V to +0.6 V, with a step resolution of 10 mV. The vOECT array is immersed in a solution containing the target marker. Under this scan sequence, the steady-state transition characteristics of the c-th independent sensor channel are recorded as a drain current vector. The acquired raw multi-channel current responses are time-synchronized, and the current response vectors of all independent sensor channels are concatenated to construct a high-dimensional signal matrix.

[0028] in, This is a multi-channel integrated input signal matrix under a single measurement. Let be the drain current vector recorded for the c-th independent sensor channel. This represents the total number of independent sensor channels in the vOECT array.

[0029] S102: Construct a feature generation module driven by a large language model, using the large language model as a semantic-aware genetic operator, performing logical-aware mutation and structural-functional crossover in the symbolic code space, evolving to generate physically interpretable feature extraction scripts for each independent sensing channel, and mapping the high-dimensional signal matrix into low-dimensional physical features.

[0030] In this embodiment, Qwen-Plus is used as the large language model. During the evolutionary stage, the generation temperature of the large language model is a key control variable for regulating the balance between genetic diversity and code syntax accuracy in the evolving population. The temperature parameter needs to be limited to a specific numerical range. If the temperature parameter is set too high (e.g., greater than 1.0), the model output will introduce excessive randomness and illusions, leading to frequent syntax errors or insecure library calls in the generated Python scripts, significantly reducing the pass rate of the isolation sandbox security verification. If the temperature parameter is set too low (e.g., less than 0.2), the model output will tend to be highly deterministic, causing mutation and crossover operations to degenerate into meaningless repetitive code generation, losing population diversity, and making the evolutionary process prone to getting trapped in local optima. In this embodiment, the temperature is set to 0.8 during the evolutionary stage. Meanwhile, because the evolutionary algorithm needs to re-inject the quantized multi-objective mixed evaluation score and the execution error stack of the isolated sandbox feedback as semantic feedback signals into the prompt words through a closed-loop evolutionary feedback mechanism, and the generated code must be an independent script containing complete defensive programming logic, the model must support a sufficiently large context receiving window and a large maximum output length limit. This ensures that the feature extraction algorithm or meta-learner evolutionary algorithm will not produce incomplete syntactic structures due to output truncation during the generation process. In this embodiment, the maximum token is set to 8000. The Python script is stored in the form of source code text and encoded into an evolvable "individual" by combining the Abstract Syntax Tree (AST) for static syntax verification and interface constraints. After the genetic operator generates or mutates to produce new code text, the evolutionary engine first calls the built-in Python ast.parse to compile and analyze the text. If a SyntaxError occurs, or if the script is found to be missing a necessary interface function signature (such as extract_features not being defined) after traversing the AST nodes, the individual is directly determined to be invalid and eliminated.

[0031] In one embodiment of this application, the feature generation module driven by the large language model performs population initialization through a domain-specific prompt word template. The prompt word template encapsulates the electrochemical physics prior knowledge of the vertical organic electrochemical transistor into structured instructions, guiding the large language model to generate an initial feature extraction script in a physically constrained symbolic space.

[0032] In one embodiment of this application, domain-specific prompt word templates are used to perform heuristic population initialization. The prompt word template system consists of three levels: system prompt words, gene evolution prompt words (including generation, mutation, and crossover prompt words), and error repair prompt words. System prompt words establish the expert role of the large language model in the evolutionary system and clarify the structural specifications of inputs and outputs, as well as the hard constraints on calling third-party computing libraries. Gene evolution prompt words input global context, including the current generation and the best performance characteristics of the current population, into the model during iteration, and provide parent generation code and performance scores to guide the model to perform local code logic mutations or code structure reorganization between multiple parent generations. Error repair prompt words are used to feed back the error stack information captured at runtime to the large language model when compilation or execution errors occur, driving it to perform targeted code defect repair. A specific example is as follows: In one embodiment of this application, encapsulating the prior knowledge of vOECT electrochemical physics into structured instructions means converting the ion-electro-coupling characteristics of the vOECT device into data structure definitions and mathematical and physical calculation specifications that can be understood by a large language model within the prompt. The specific encapsulation process includes: defining two basic physical quantities, voltage (gate voltage array) and current (drain current array), in the input interface; and explicitly listing the maximum transconductance in the prompt. Maximum transconductance voltage Threshold voltage Starting voltage Electrochemical characterization indicators such as on / off ratio serve as guiding directions for feature generation. Finally, specific solution algorithm specifications are provided to guide the large language model in transforming the solution of complex physical models into standard mathematical calculation steps such as differential solutions and univariate linear regression. Taking transconductance as an example, the instructions transform the ion implantation regulation mechanism and macroscopic electrical response law of vOECT into heuristic objectives, requiring the large language model to use the numerical central difference method to calculate transconductance and extract the maximum transconductance and its corresponding gate voltage.

[0033] Taking transconductance as an example, this embodiment illustrates how physical prior knowledge is encapsulated into structured instructions: First, the experimental measurement parameters are mapped one-to-one with the variables in the code interface within the prompt. In the instruction design, the gate scan voltage (in V) and drain current (in A) collected from the vOECT transfer characteristic curve are explicitly defined as one-dimensional numerical vectors voltage and current in the feature extraction function interface, respectively. Second, the instructions need to transform the ion implantation control mechanism and macroscopic electrical response law of vOECT into heuristic targets guiding the large model search. The instructions explicitly state that the first derivative of the drain current with respect to the gate voltage (i.e., transconductance) The transconductance curve characterizes the efficiency of gate voltage in regulating channel current, and shifts when the concentration of the external target ion changes. Therefore, the instruction requires that the large language model, when writing the Python feature extraction script, must locate and extract two core indicators: one is the maximum value of the transconductance sequence (…). ), used to characterize the maximum amplification factor of the device; and secondly, the gate voltage corresponding to the maximum transconductance ( This is used to quantitatively characterize the shift of the inflection point of the transition curve on the voltage axis. Finally, the instructions must encapsulate the rigor and operational safety of numerical calculations into specific code writing constraints. The instructions specifically specify the solution specifications for differential gradients, requiring large language models to use the numerical central difference method (e.g., calling `np.gradient(current, voltage)`) to reduce single-point measurement noise, and explicitly prohibiting the use of local single-point slope calculations that may cause abnormal fluctuations. Simultaneously, the instructions impose hard constraints on code robustness, requiring that the generated function body must include code structure checks for illegal input cases (e.g., data points less than or equal to 1, containing NaN or infinite values). In this way, a structured instruction encapsulation is completed, transforming "physical prior knowledge" into "executable code specifications that conform to numerical stability."

[0034] In one embodiment of this application, the logically-aware variation includes: The code and performance evaluation score of the selected well-performing and runnable parent feature extraction script are input into the large language model as semantic feedback context. The large language model then extends the parent feature extraction logic in mathematical or physical dimensions to generate a logically optimized offspring mutation script.

[0035] In one embodiment of this application, in each generation of evolution, the system first filters out well-performing and runnable parent feature extraction scripts through security verification and a multi-objective hybrid evaluation function. When the evolution engine triggers the mutation operator, the complete code of the selected parent script and its performance evaluation score are encapsulated in a prompt word as semantic feedback context and sent to the large language model. The large language model leverages its code understanding and logical reasoning capabilities to extend the parent's feature extraction logic in mathematical or physical dimensions, such as adding new feature calculations, improving the calculation accuracy of existing features, introducing more complex nonlinear transformations, or exploring unexplored combinations of physical parameters, thereby generating logically superior offspring mutant individuals. Erroneous code is eliminated during the evaluation phase and will not participate in the mutation operation as a parent.

[0036] In one embodiment of this application, the structural functional intersection includes: The code of the two parent feature extraction scripts is simultaneously input into the large language model. The large language model analyzes and reorganizes the feature extraction logic of the two parent scripts at the semantic level, retains complementary computational operators, removes redundant features, resolves naming conflicts and function dependencies, and generates a child cross script that integrates advantageous logic.

[0037] In one embodiment of this application, the system simultaneously inputs the codes of parent A and parent B into a large language model, guiding it to analyze and reorganize the feature extraction logic of the two parents at the semantic level. The large language model identifies the complementarity and correlation between the extracted features of the two parents, retains computational operators with complementary information content and different physical meanings in the two sets of features, eliminates redundant features with high linear correlation, and adaptively resolves variable naming conflicts and functional dependencies during the merging process. Ultimately, it organically integrates the advantageous computational logic of both into a single, structurally complete child script, achieving deep cross-functionality at the semantic level. Because evolution is not completed in one generation but through multiple generations, a single feature extraction script is finally obtained. The first generation's parent A contains features ABC, and the second generation evolves features CDE. The two structures and functions cross-reference, resulting in parent C containing features ABCDE. The extraction method for C may differ in different parents. The large model selects the superior extraction method based on the score and removes the other extraction method, ultimately completing the structural-functional cross-reference between the two parents.

[0038] Through the iterative evolution of logical perception variation and structural function intersection, a physically interpretable feature extraction script for each independent sensing channel is finally generated, mapping the high-dimensional signal matrix into low-dimensional physical features.

[0039] S103: Construct a meta-learner optimization module driven by a large language model, and use the large language model to evolve and generate a meta-learner, which is used to adaptively allocate multi-channel decision weights.

[0040] In one embodiment of this application, the large language model-driven meta-learner optimization module further includes: The low-dimensional physical features are input into the base learner to generate concentration decoupling values ​​for each independent sensing channel, and the concentration decoupling values ​​of each channel are concatenated to construct a decoupling matrix. The large language model is used to perform meta-learner evolution in a preset regression algorithm cluster space. The evolution includes the evolution of algorithm type and the evolution of algorithm hyperparameters. The meta-learner takes the decoupling matrix as input and outputs the global concentration decoupling value after allocating decision weights for each channel.

[0041] In one embodiment of this application, low-dimensional physical features are input into the base learner to generate independent channel concentration decoupling values. The algorithm structure and hyperparameters of the base learner are not a unified fixed model, but are determined by the large language model through autonomous evolution and search within a pre-defined regression algorithm cluster space (including linear regression, regularized models, Gaussian process regression, support vector regression, random forest, and Huber regression). In the first stage of evolution, the large language model selects and optimizes within the algorithm space based on the electrical response characteristics of each specific channel, generating corresponding function scripts. Therefore, the base learners for different channels obtained in the final evolution may have heterogeneous and different algorithm types and hyperparameter configurations, depending on the sensitivity of each channel's data response to different models.

[0042] The large language model is used to perform meta-learner evolution, searching within a predefined algorithm family space. Specifically, the algorithm space that the large model is allowed to search and call is based on standard open-source machine learning libraries (such as scikit-learn). The meta-learner evolution range includes: weighted average strategies, regularized linear regression models (Ridge, Lasso, HuberRegressor), and nonlinear ensemble models (SVR, RF, GradientBoostingRegressor). The large language model evolves the meta-learner in two dimensions—algorithm type and algorithm hyperparameters—by rewriting the code: the evolution of the algorithm type determines the level of fusion logic used (e.g., evolving from weighted average to linear regression and then to support vector machine regression); the evolution of the algorithm hyperparameters is directly modified in the generated Python code (e.g., adjusting the regularization coefficient alpha in ridge regression). The meta-learner takes the decoupling matrix as input and outputs the global concentration decoupling value after assigning decision weights to each channel. It should be noted that the underlying closed-loop control mechanism of LLM in the meta-learner evolution is exactly the same as that of the feature extraction script. Both methods employ large models as semantic genetic operators, undergoing an iterative process of "template generation --> LLM code rewriting and local reconstruction --> AST static parsing and syntax filtering --> sandbox data fitting and validation set scoring --> optimal program selection". They differ in some aspects, primarily in the input: the feature extraction script's evolution input is the original physical sequence data, while the meta-learner's evolution input is a concentration decoupling matrix composed of base learner concentration decoupling values; the feature extraction script's evolution is constrained by electrochemical prior knowledge, while the meta-learner's evolution is not.

[0043] In one embodiment of this application, the base learners of each independent sensing channel are trained independently and do not share parameters. The algorithm type and hyperparameter configuration of each base learner are autonomously evolved by the large language model for the electrical response characteristics of the corresponding channel.

[0044] In one embodiment of this application, regarding the training method of the base learner, in the first stage, the feature extraction script and the base learner model are in a state of joint evaluation and co-optimization. For each candidate generated by the large language model, the evolutionary engine first extracts channel features, then uses these features to train the base learner, and evaluates and iteratively filters the candidate through the evaluation function score. After the first stage, the optimal feature extraction script and the code structure of the base learner for each channel are fixed. In the second stage, each channel's base learner only uses the parameters fitted in the first stage to perform forward prediction. Its parameters are no longer jointly optimized with the meta-learner in the second stage. The meta-learner evolves and is trained independently based on the fixed prediction output of the base learner. The base learner for each independent channel is trained completely independently and does not share parameters. Due to the random heterogeneity in the fabrication process of each channel device in the vOECT array, there are certain differences in the electrical characteristics and baseline drift of each channel. Therefore, each channel independently evolves its feature extractor and base learner, and only uses the drain current-gate voltage curve data collected from that specific channel to train the corresponding base learner.

[0045] S104: Deploy the evolved feature extraction script and the meta-learner in an isolated computing sandbox environment, perform security verification and quantification scoring by combining a multi-objective hybrid evaluation function, and drive the feature extraction script and the meta-learner to converge toward the optimal signal decoupling model through closed-loop feedback.

[0046] In one embodiment of this application, the candidate feature extraction scripts and meta-learners are dynamically compiled within a multi-process isolation sandbox equipped with computational resource limits and timeout interrupt mechanisms. The sandbox intercepts illegal calls and resource overruns, ensuring the safe operation of the evolutionary engine.

[0047] After feature extraction is completed within the isolated process, the system calls the trained base learner to perform preliminary validation, generating independent channel concentration decoupling values. Simultaneously, non-negative truncation is introduced at the model output layer.

[0048] in, The original value decoded by the algorithm. This is the final output concentration after truncation. This non-negative truncation process forcibly corrects negative values ​​in the original values ​​of the regression model output to zero, ensuring that the output concentration meets the physical benchmarks for electrochemical detection.

[0049] In one embodiment of this application, the multi-objective hybrid evaluation function comprehensively measures the mean absolute error, mean absolute percentage error, coefficient of determination, and cross-channel stability index, and outputs a comprehensive evaluation score as the sole metric guiding the closed-loop feedback.

[0050] In one embodiment of this application, a multi-objective hybrid fitness function is used for quantitative scoring. In the repeated random partitioning of the cross-validation stream, the comprehensive Selection Score is calculated as the sole metric guiding the closed-loop feedback:

[0051]

[0052] in, The comprehensive evaluation score used to guide the closed-loop feedback of the evolution engine; , and These are the mean absolute error, mean absolute percentage error, and mean coefficient of determination calculated in K repeated random partitioning assessments, respectively. and These are the standard deviations of the corresponding indicators in multiple evaluations, used to characterize the numerical stability of the model; , , as well as These are the preset weighting coefficients for each component; where, and This is the preset error normalization scaling factor; This is a preset stability penalty normalization scaling factor. The scaling factor is used to eliminate the differences in dimensions and orders of magnitude among the evaluation indicators, so that each indicator is mapped to a similar numerical range, thereby ensuring that the weight coefficients can effectively control the multi-objective optimization preferences of the evolution engine.

[0053] The specific formulas for calculating the mean absolute error, mean absolute percentage error, and coefficient of determination are as follows:

[0054]

[0055]

[0056] in, For the first The true biomarker concentration of each sample; The decoupling concentration is the original output of the model; The non-negative truncation operator is applied to the predicted concentration to ensure that the decoupling results conform to physical reality; N is the total number of samples in a single evaluation. This is the number of valid samples after removing zero overflow and non-finite values. This is the arithmetic mean of the true concentrations.

[0057] Cross-channel stability penalty The calculation method is as follows:

[0058] in, This represents the total number of repeated evaluations. For the first The mean absolute error obtained under the subdivision evaluation.

[0059] A closed-loop evolutionary feedback mechanism is constructed to convert the quantified comprehensive evaluation score into a semantic feedback signal and re-inject it into the evolutionary engine driven by the large language model. This selects excellent features and drives the feature extraction script and meta-learner to converge towards the globally optimal signal decoupled model through iterative optimization.

[0060] S105: Use the trained optimal signal decoupling model to decouple the multi-channel response sequence of the organic sensor array and output a high-fidelity decoupling result of the target marker concentration.

[0061] In this embodiment, the raw multi-channel current response sequence acquired by a multi-channel vOECT array deployed in an actual monitoring scenario is obtained, preprocessed, and a high-dimensional signal matrix is ​​constructed. The high-dimensional signal matrix is ​​then input into a trained optimal signal decoupling model. First, the model's globally optimal feature extraction script extracts feature vectors reflecting the transient response dynamics of the device. Subsequently, these feature vectors are input into multiple parallel-constructed base learners, which output the decoupling values ​​of the marker concentrations for each independent channel. A decoupling matrix composed of the decoupling values ​​of the marker concentrations for each channel is constructed and input into the optimal meta-learner. The meta-learner dynamically adjusts the decision weights of each channel to neutralize the bias caused by the nonlinear heterogeneity of the device, ultimately outputting a high-fidelity global decoupling result.

[0062] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0063] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 2As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a high-fidelity information decoupling method for organic sensor arrays driven by a large model. The display screen can be an LCD screen or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0064] Those skilled in the art will understand that Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0065] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0066] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0067] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0068] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0069] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0070] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0071] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution, characterized in that, The method includes: Obtain the multi-channel sensor response sequence of the organic sensor array and construct a high-dimensional signal matrix; A feature generation module driven by a large language model is constructed. The large language model is used as a semantic-aware genetic operator. Logical-aware mutation and structural-functional crossover are performed in the symbolic code space to evolve and generate physically interpretable feature extraction scripts for each independent sensing channel, and the high-dimensional signal matrix is ​​mapped to low-dimensional physical features. A large language model-driven meta-learner optimization module is constructed, which uses the large language model to evolve and generate a meta-learner. The meta-learner is used to adaptively allocate multi-channel decision weights. The feature extraction script and the meta-learner generated by evolution are deployed in an isolated computing sandbox environment. Security verification and quantitative scoring are performed by combining a multi-objective hybrid evaluation function. The feature extraction script and the meta-learner are driven to converge toward the optimal signal decoupling model through closed-loop feedback. The optimal signal decoupling model, after training, is used to decouple the multi-channel response sequence of the organic sensor array, and outputs a high-fidelity decoupling result of the target marker concentration.

2. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The organic sensing array is a vertical organic electrochemical transistor array.

3. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The feature generation module driven by the large language model performs population initialization through domain-specific prompt word templates. These prompt word templates encapsulate the electrochemical physics prior knowledge of vertical organic electrochemical transistors into structured instructions, guiding the large language model to generate an initial feature extraction script within a physically constrained symbolic space.

4. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The logical perception variation includes: The code and performance evaluation score of the selected well-performing and runnable parent feature extraction script are input into the large language model as semantic feedback context. The large language model then extends the parent feature extraction logic in mathematical or physical dimensions to generate a logically optimized offspring mutation script.

5. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The structural functional intersection includes: The code of the two parent feature extraction scripts is simultaneously input into the large language model. The large language model analyzes and reorganizes the feature extraction logic of the two parent scripts at the semantic level, retains complementary computational operators, removes redundant features, resolves naming conflicts and function dependencies, and generates a child cross script that integrates advantageous logic.

6. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The large language model-driven meta-learner optimization module further includes: The low-dimensional physical features are input into the base learner to generate concentration decoupling values ​​for each independent sensing channel, and the concentration decoupling values ​​of each channel are concatenated to construct a decoupling matrix. The large language model is used to perform meta-learner evolution in a preset regression algorithm cluster space. The evolution includes the evolution of algorithm type and the evolution of algorithm hyperparameters. The meta-learner takes the decoupling matrix as input and outputs the global concentration decoupling value after allocating decision weights for each channel.

7. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 6, characterized in that, The base learners for each independent sensing channel are trained independently and do not share parameters. The algorithm type and hyperparameter configuration of each base learner are autonomously evolved by the large language model for the electrical response characteristics of the corresponding channel.

8. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The isolated computing sandbox is a multi-process isolated sandbox with computing resource limits and timeout interruption mechanisms. It is used to dynamically compile the feature extraction script and the meta-learner, intercept illegal calls and resource overruns, and ensure the safe operation of the evolution engine.

9. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The multi-objective hybrid evaluation function comprehensively measures the mean absolute error, mean absolute percentage error, coefficient of determination, and cross-channel stability index, and outputs a comprehensive evaluation score as the sole metric guiding the closed-loop feedback.

10. The high-fidelity information decoupling method for organic sensor arrays driven by large-scale model evolution according to claim 1, characterized in that, The method further includes: A non-negative truncation process is introduced at the model output to force the negative values ​​in the original values ​​of the regression model output to be corrected to zero, ensuring that the output concentration meets the physical benchmark of electrochemical detection.