Cost prediction method and device for biological sample library and computing equipment
By analyzing biobank cost data using multi-task neural networks and target residual networks, the problem of low cost accounting efficiency in biobanks is solved, enabling rapid and accurate cost prediction and meeting the cost management needs of high-throughput, multi-process scenarios.
Patent Information
- Application Number
- CN202511490066.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-12-12
AI Technical Summary
Existing technologies for biobank cost accounting are inefficient and fail to meet the cost management needs of high-throughput, multi-process scenarios.
By acquiring initial cost data to generate a cost vector, using a multi-task neural network to analyze time series and environmental characteristics, and combining the target residual network to allocate indirect costs, a total cost prediction value for each biological sample is generated.
It enables rapid and accurate prediction of the cost of each biological sample, improves cost accounting efficiency, and meets the needs of efficient management of biobanks.
Smart Images

Figure CN121119291A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to a method, apparatus, computing device, and computer program product for predicting the cost of a biobank. Background Technology
[0002] With the development of precision medicine, biobanks have evolved from "storage centers" to "data-cost centers." This transformation has placed unprecedented demands on the operation and management of biobanks, especially cost management. Currently, biobank cost management mainly relies on Hospital Information Systems (HIS), Laboratory Information Management Systems (LIMS), and Enterprise Resource Planning (ERP) systems. These systems enable financial recording at the institutional or departmental level. To improve cost accounting accuracy, traditional Activity-Based Costing (ABC) and its evolution, Time-Driven Activity-Based Costing (TDABC), have been introduced. However, in the high-throughput, multi-process scenario of modern biobanks, these methods heavily rely on manually pre-defining and maintaining complex operational processes and driver parameters. Furthermore, the models often face the challenge of "dimensionality explosion," leading to high human maintenance costs and low accounting efficiency, making it difficult to meet the urgent needs of biobank cost management. Summary of the Invention
[0003] This application provides a method, apparatus, computing device, and computer program product for predicting the cost of biobanks, in order to solve the problem of low cost accounting efficiency in the prior art.
[0004] In a first aspect, embodiments of this application provide a method for cost prediction of a biobank, comprising: acquiring initial cost data; generating a cost vector based on the initial cost data, wherein the initial cost data represents the expenses incurred in processing biological samples; generating an input vector based on the cost vector; analyzing the input vector using a multi-task neural network to obtain a first cost prediction value, wherein the multi-task neural network is used at least to analyze the temporal and environmental characteristics of the input vector, the environmental characteristics characterizing the operating environment of the biological samples, the first cost prediction value including at least direct costs and indirect costs, the direct costs representing the expenses directly incurred in processing each biological sample, and the indirect costs representing the expenses incurred in processing multiple biological samples; analyzing at least the first cost prediction value using a target residual network to obtain a weight matrix, wherein the target residual network is used to analyze the causes of the cost value of each biological sample, and the weight matrix represents the weight of the indirect costs allocated to each biological sample; allocating the indirect costs using the weight matrix to obtain a second cost prediction value for each biological sample; and calculating a total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value.
[0005] Secondly, this application provides a cost prediction device for a biobank, comprising: A generation unit is used to acquire initial cost data and generate a cost vector based on the initial cost data, wherein the initial cost data represents the costs incurred in processing biological samples; The first analysis unit is used to generate an input vector based on the cost vector, and analyze the input vector through a multi-task neural network to obtain a first cost prediction value. The multi-task neural network is used to analyze at least the temporal and environmental features of the input vector. The environmental features characterize the operating environment of the biological sample. The first cost prediction value includes at least direct costs and indirect costs. The direct costs represent the expenses directly incurred in processing each biological sample, and the indirect costs represent the expenses incurred in processing multiple biological samples. The second analysis unit is used to analyze at least the first cost prediction value through a target residual network to obtain a weight matrix, wherein the target residual network is used to analyze the causes of the cost value of each biological sample, and the weight matrix represents the weight of the indirect cost allocated to each biological sample. The allocation unit is used to allocate the indirect costs through the weight matrix to obtain a second cost prediction value for each biological sample; The calculation unit is used to calculate the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value.
[0006] Thirdly, this application provides a computing device, including a processing component and a storage component; the storage component stores a computer program; the computer program is invoked and executed by the processing component to implement any of the aforementioned cost prediction methods for biobanks.
[0007] Fourthly, this application provides a computer program product, including a computer program or instructions, which, when executed by a processing component, implement any of the aforementioned cost prediction methods for biobanks.
[0008] This application embodiment obtains initial cost data and generates a cost vector based on it. An input vector is generated from the cost vector, and a multi-task neural network is used to analyze the input vector to obtain a first predicted cost value. This first predicted cost value includes at least direct and indirect costs. A target residual network is then used to analyze at least the first predicted cost value to obtain a weight matrix. The indirect costs are then allocated using the weight matrix to obtain a second predicted cost value for each biological sample. Finally, the total predicted cost value for each biological sample is calculated based on the first and second predicted cost values. By separately predicting the direct and indirect costs of each biological sample using a multi-task neural network and a target residual network, and allocating the indirect costs, the cost of each biological sample can be predicted quickly. Therefore, this addresses the problem of low cost accounting efficiency in existing biobank technologies, achieving the goal of improving cost accounting efficiency.
[0009] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This paper illustrates a system architecture diagram in which the technical solutions of the embodiments of this application can be applied; Figure 2 A flowchart illustrating a method for predicting the cost of a biobank according to an embodiment of this application is shown. Figure 3 This illustration shows a structural schematic diagram of a multi-source cost-aware engine provided in an embodiment of this application; Figure 4 This illustration shows a schematic diagram of the structure of a multi-task neural network provided in an embodiment of this application; Figure 5 This illustration shows a schematic diagram of an indirect cost attribution and allocation process provided in an embodiment of this application; Figure 6 This illustration shows a schematic diagram of a model closed-loop learning and training process provided in an embodiment of this application; Figure 7 This application illustrates a flowchart of a federal compliance audit chain provided in an embodiment. Figure 8 This paper shows a schematic diagram of the structure of a cost prediction device for a biobank provided in an embodiment of this application; Figure 9 This illustration shows a structural schematic diagram of one embodiment of a computing device provided in this application. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0012] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0013] Additionally, it should be noted that when user interaction operations or triggering operations are involved in the embodiments of this application, these operations include, but are not limited to, various interaction methods such as touch operations, gesture operations, voice operations, head movement operations, and eye movement operations. Touch operations include, but are not limited to, click operations, double-click operations, long-press operations, swipe operations, pinch operations, or mouse hover operations. Swipe operations include, but are not limited to, straight-line swipes and curved-line swipes.
[0014] Furthermore, it should be noted that when the embodiments of this application involve jumping between the first interface and the second interface, the jumping methods involved in the embodiments of this application include, but are not limited to: jumping directly from the first interface to the second interface, or jumping from the first interface to the task interface and completing the corresponding task operation on the task interface before jumping to the second interface; completing the corresponding task operation on the task interface includes, but is not limited to: completing the game operation on the game interface when the task interface is implemented as a game interface; completing identity authentication on the identity authentication interface when the task interface is implemented as an identity authentication interface; and completing the recharge operation on the recharge interface when the task interface is implemented as a recharge interface.
[0015] To address the problem of low cost accounting efficiency in existing biobank technologies, this application provides a solution. The basic idea is as follows: Initial cost data is acquired, and a cost vector is generated based on this data. An input vector is generated from the cost vector, and a multi-task neural network is used to analyze the input vector to obtain a first predicted cost value. This first predicted cost value includes at least direct and indirect costs. A target residual network is then used to analyze at least the first predicted cost value to obtain a weight matrix, and the indirect costs are allocated using this weight matrix to obtain a second predicted cost value for each biosample. The total predicted cost value for each biosample is calculated based on the first and second predicted cost values. By using both the multi-task neural network and the target residual network to predict the direct and indirect costs of each biosample and then allocating the indirect costs, the cost of each biosample can be predicted quickly. Therefore, this solution addresses the problem of low cost accounting efficiency in existing biobank technologies, achieving the goal of improving cost accounting efficiency.
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] Figure 1 A system architecture diagram of a technical solution of an embodiment of this application is shown, which can be applied thereto. The system architecture may include a user terminal 101 and a server terminal 102.
[0018] In this system, the user terminal 101 and the server terminal 102 can establish a connection via a network. The network provides a communication link between the user terminal 101 and the server terminal 102. The network can include various connection types, such as wired, wireless, or fiber optic cables. The user terminal 101 can interact with the server terminal 102 through the network to receive or send messages, etc.
[0019] The user terminal 101 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The user terminal 101 can be deployed on electronic devices and depends on the device or certain apps on the device to run. Electronic devices can have displays and support information browsing, such as personal mobile terminals like mobile phones, tablets, personal computers, desktop computers, smart speakers, smartwatches, etc. For ease of understanding... Figure 1 The user end is primarily represented by the image of a device. Various other types of applications can also be configured in electronic devices, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software. Electronic devices can refer to devices used by users that have the computing, internet access, and communication functions required by the user, such as mobile phones, tablets, personal computers, and wearable devices. Electronic devices typically include at least one processing component and at least one storage component. Electronic devices may also include basic configurations such as network interface cards (NICs), I / O (input / output) buses, and audio / video components; this application does not limit their inclusion of these components. Optionally, depending on the implementation of the electronic device, it may also include some peripheral devices, such as keyboards, mice, input pens, and printers; this application does not limit their inclusion of these components.
[0020] Server 102 may include servers that provide various services, such as a server for backend training that supports the model used on client 101, or a server that processes interactive information sent by client.
[0021] It should be noted that server 102 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0022] It should be noted that the biobank cost prediction method provided in this application embodiment is generally executed by the server 102, and the corresponding biobank cost prediction device is generally located in the server 102. However, in other embodiments of this application, the user terminal 101 may also have similar functions to the server 102, thereby executing the biobank cost prediction method provided in this application embodiment. In other embodiments, the biobank cost prediction method provided in this application embodiment may also be jointly executed by the user terminal 101 and the server 102. It should be understood that Figure 1 The number of client and server instances shown is merely illustrative. Depending on implementation needs, there can be any number of client and server instances.
[0023] The implementation details of the technical solutions in the embodiments of this application are described in detail below.
[0024] Figure 2 This document illustrates a flowchart of a method for predicting the cost of a biobank, as provided in an embodiment of this application. Figure 2 As shown, it includes the following steps: S201, Obtain initial cost data, and generate a cost vector based on the initial cost data, wherein the initial cost data represents the expenses incurred in processing biological samples; In one embodiment of this application, the initial cost data can be cost data from processes such as testing, storage, transportation, quality control, manpower, consumables, and depreciation of biological samples. This data is collected through a multi-source cost-aware engine, and the collected data undergoes real-time cross-system acquisition and heterogeneous alignment to output a cost vector with uniform granularity. In practical applications, the initial cost data is not limited to data from the aforementioned processes and can be cost data from any actual operational process involving biological samples.
[0025] S202, an input vector is generated based on the cost vector, and the input vector is analyzed by a multi-task neural network to obtain a first cost prediction value. The multi-task neural network is used to analyze at least the temporal and environmental features of the input vector. The environmental features characterize the operating environment of the biological sample. The first cost prediction value includes at least direct costs and indirect costs. The direct costs represent the expenses directly incurred in processing each biological sample, and the indirect costs represent the expenses incurred in processing multiple biological samples. In one embodiment of this application, a multi-task neural network architecture integrating temporal-graph neural networks is employed. Using the sample lifespan as the axis, it predicts the first cost of a single biological sample. The temporal-graph neural network can analyze the time-series features and environmental characteristics of the input vector. The temporal coding branch uses an Informer encoder to capture long-range dependencies in the cost sequence; the graph convolution branch models indirect cost propagation using a sample-device-person hypergraph; and the multi-task neural network simultaneously outputs direct cost, indirect cost, and QALY (quality-adjusted life-year contribution).
[0026] S203, at least the first cost prediction value is analyzed through the target residual network to obtain a weight matrix, wherein the target residual network is used to analyze the cause of the cost value of each biological sample, and the weight matrix represents the weight of the indirect cost allocated to each biological sample. In one embodiment of this application, cost drivers are automatically attributed based on a target residual network, and public costs are dynamically allocated to individual biological samples with interpretable weights. The target residual network can be a dual residual network.
[0027] S204, the indirect costs are allocated using the weight matrix to obtain a second cost prediction value for each biological sample. In one embodiment of this application, the above-mentioned dual residual network is used to generate an allocation weight matrix W that satisfies ∑W_ij=1. The interpretability constraint Grad-cam is introduced to ensure that the allocation logic meets the readability requirements of management audit, and a second cost prediction value for each of the above-mentioned biological samples is obtained, that is, the cost of indirect costs allocated to each biological sample.
[0028] S205, the total cost prediction value for each of the above-mentioned biological samples is calculated based on the first cost prediction value and the second cost prediction value.
[0029] In one embodiment of this application, the total cost prediction value for each biological sample is obtained by adding the direct and indirect costs allocated to each biological sample in the first cost prediction value (the second cost prediction value).
[0030] In some optional implementations, step S201 can be achieved through the following steps: Step S2011: Obtain structured cost data, parse and store the structured cost data using a protocol adaptation template to obtain first initial cost data, wherein the structured cost data follows data storage rules; Step S2012: Obtain unstructured cost data, perform image processing on the unstructured cost data to convert it into structured cost data, obtaining second initial cost data; Step S2013: Concatenate the first initial cost data and the second initial cost data to obtain the cost vector. This method, through the above steps, can convert initial cost data of various structures, accelerating data conversion efficiency and facilitating subsequent cost prediction and calculation using a network model based on the initial cost data.
[0031] In one embodiment of this application, such as Figure 3 The diagram shows a schematic of the multi-source cost-aware engine. The multi-source cost-aware engine is used to obtain initial cost data. First, the system topology of the multi-source cost-aware engine includes: (1) Edge layer: Each sample library ultra-low temperature freezer, automated liquid nitrogen tank, and transport cold chain box is equipped with a lightweight MQTT / OPC-UA network 5173 (ARM Cortex-A72 quad-core, 2 GB RAM). (2) Middle layer: K3s lightweight Kubernetes cluster (3 nodes, 16 vCPU / 64 GB RAM) is responsible for protocol parsing, caching, and streaming ETL. (3) Core layer: Distributed time-series database TDengine3.0 (three replicas), with a daily write volume of 120 million records and a compression ratio of 8:1; hot data is retained for 30 days, and cold data is automatically pushed down to MinIO object storage (erasure coding 10+2).
[0032] Secondly, for structured cost data, the multi-source cost-aware engine includes a protocol adapter (micro-service). As a front-end micro-service of the multi-source cost-aware engine, the protocol adapter uses a "one-time parsing, two-level caching, semantic mapping" mechanism to perform sub-second parsing, field-level alignment, and atomic event output on the edge side for data packets from three heterogeneous protocols: HL7 FHIR, DICOM SR, and IHE LAB LTW. The protocol conversion rules are shown in Table 1, thereby simultaneously eliminating protocol conversion latency and improving cross-system data alignment accuracy. Specifically: Pre-built protocol templates and field-level semantic dictionaries: An XSD / Schema template is established for each protocol, and the field path corresponding to the "Unified Event Code (UEC)" is marked in the template; SNOMED CT and LOINC standard code tables are introduced as semantic dictionaries to achieve automatic unification of synonymous fields (e.g., DICOM "StudyInstance UID"). FHIR ("imagingStudy.reference"). Zero-copy streaming parsing: Using the Netty zero-copy framework, TCP / MLLP packets are sliced directly in the kernel buffer after being received, avoiding the ≥200 ms I / O latency caused by the traditional "write to disk first and then parse" method; the lightweight StAX parser of HAPI-FHIR is used for HL7 messages, and Dcm4che's streaming Part-10 reading is used for DICOM messages, with a single message parsing time of <15 ms (P99). The two-level caching and incremental timestamp alignment include: L1 edge cache (32 MB circular buffer): storing the raw messages from the most recent 5 minutes for out-of-order reordering; L2 in-memory database (Redis with RediSearch): using "SampleID + UTC seconds" as the composite key, caching parsed atomic events and supporting cross-system join queries within <5 ms; before outputting atomic events, the parsing thread first checks whether the same "SampleID + UTC seconds" key-value already exists in L2: if it exists, the amount and attribute are merged; if it does not exist, a new key is written, ensuring accurate alignment of multi-protocol events within the same second-level time window. Adaptive backpressure and circuit breaker mechanism: When the QPS of a single node exceeds the threshold (FHIR 2000, DICOM 1500, IHE LAB 1200), backpressure is automatically triggered, and subsequent packets are temporarily stored in Kafka to avoid latency avalanche caused by parsing blocking. If the parsing thread throws a verification exception three times consecutively, the protocol adapter is circuit-broken for 30 seconds, during which traffic is automatically routed to a backup instance to ensure overall throughput. Through the above methods, the protocol adapter completes the real-time conversion of "protocol → semantics → atomic events" at the edge, with a parsing latency P99 < 50 ms and a cross-system data alignment accuracy ≥ 99.2%, thereby directly solving the two major technical problems of "protocol conversion latency" and "insufficient cross-system data alignment accuracy". As a front-end microservice of the "multi-source cost-aware engine", the protocol adapter uses a "one-time parsing, two-level caching, semantic mapping" mechanism to complete sub-second parsing, field-level alignment, and atomic event output of data packets of three heterogeneous protocols, namely HL7 FHIR, DICOM SR, and IHE LAB LTW, at the edge side, thereby eliminating protocol conversion latency and improving cross-system data alignment accuracy.
[0033] Table 1 Protocol Conversion Rules For unstructured cost data, such as invoice recognition units, the following methods are used: Obtain the original unstructured invoice file (PDF / A-2b, TIFF G4, JPEG XL format, maximum 50MB), and receive the uploaded invoice file through a file gateway. Key fields are identified using a fine-tuned LayoutLMv3, and the model further processes the invoice image and text (if present), identifying 21 predefined key fields (such as invoice code, number, amount, tax rate, reagent batch number, and specifications). The model outputs the original recognition results for these fields. Finally, structured output and validation are performed: a CRF sequence labeling layer optimizes field sequence relationships, and regular expressions correct the format and perform logical validation (such as checksum verification) on the recognition results (especially key fields like amount and code). The amount field requires 100% validation success. The final output is a structured invoice information record (JSON or database record) containing the identified fields and their values. Flow: As part of the original cost data, it is input into the engine's unified processing flow, ultimately forming a cost vector. The measured character accuracy is 98.7%.
[0034] To detect abnormal data in a timely manner, a real-time anomaly detection module is also set up. This module receives real-time or near-real-time fee data streams from the aforementioned protocol adapter and after ticket recognition, performs feature extraction and encoding, and a variational autoencoder (VAE) receives the fee data stream (organized by time window). The encoder (3-layer 1D-CNN) compresses the input into a 128-dimensional latent variable z.
[0035] Reconstruction and Error Calculation: The decoder (symmetric deconvolution) reconstructs the input data based on z. The reconstruction error (the difference between the original input and the reconstructed output) is calculated.
[0036] Anomaly detection: Calculate the distribution of reconstruction error (based on a 7-day sliding window) and set a dynamic threshold (e.g., P99.5). If the reconstruction error of the current sample exceeds the threshold, it is considered an anomaly. Use β-VAE (β=4.0) + KL-annealing for training to balance reconstruction accuracy and latent variable distribution.
[0037] Output data: Abnormal alarm events (including abnormal data points, timestamps, related sample / device IDs, reconstruction error values, and possible cause predictions). Published via the Kafka topic `cost_anomaly`, it connects to downstream alarm systems (WeChat / DingTalk / Email) and implements alarm storm suppression. F1 ≥ 0.92, P99 latency < 28ms.
[0038] In some optional implementations, step S202, which generates an input vector based on the cost vector, can be achieved through the following steps: Step S2021: Analyze the cost vector using a time-series analysis network to obtain a time-series feature vector, where the time-series feature vector represents the time-series characteristics of the cost vector; Step S2022: Obtain the attribute data of the biological sample, and generate an attribute feature vector based on the attribute data, where the attribute data at least includes sample type, and the attribute feature vector represents the attribute characteristics of the biological sample; Step S2023: Obtain the usage information of the biological sample, construct a graph database model based on the usage information, and analyze the usage information of the biological sample based on the graph database model to obtain a graph feature vector, where the usage information at least includes the usage environment and users of the biological sample, and the graph feature vector represents the environmental characteristics of the biological sample; Step S2024: Concatenate the time-series feature vector, the attribute feature vector, and the graph feature vector to obtain the input vector. This method further generates the model's input vector based on the cost vector, thus extracting features from multiple data sources and more comprehensively predicting the cost of the biological sample.
[0039] In one embodiment of this application, the cost vector is the cost vector generated in the above steps. Attribute data is static attribute data of biological samples (e.g., ICD-10 encoding [multi-label Top3], sample type [One-Hot 22-dimensional], acquisition method [Embedding 8-dimensional]), which is relational data used to describe the sample bank operating environment.
[0040] The generation process of the time-series feature vector: Based on the cost vector, Apache Flink CEP is used to identify key events (such as refrigerator door opening and closing), calculate the statistical features (mean, variance, peak count) of a 30-minute sliding window, and output the time-series feature vector (Parquet format). The generation process of the graph feature vector: Hypergraph construction: A hypergraph containing four types of nodes—Sample, Device, Person, and ReagentLot—is constructed in the Neo4j graph database. Relationships (edges) include attributes: use_time (min), temp (°C), humidity (%), and event_ts. For example, a sample operation will create a hyperedge connecting the sample node, the device node, the operator node, and the reagent batch number node, recording the environmental and resource consumption information of this operation. The time-series feature vector, attribute feature vector, and graph feature vector extracted from the hypergraph (which will be learned by the model's graph branch later) are integrated to form the complete input vector of the model.
[0041] In practical applications, the network structure for temporal feature vectors can be an Informer network structure: ProbSparse Attention top-k=5; Distilling layer compression ratio 4×; Location encoding: Time2Vec (k=64). Temporal branch processing: The Informer encoder (using ProbSparse Attention, Distilling Layer) processes the sample-related cost temporal features, captures long-range dependencies, and outputs temporal feature encodings. Time2Vec location encoding is used.
[0042] The graph feature vectors are processed using a hypergraph convolution branch: HyperGCN depth L=3; hidden layer dimensions 256→128→64; DropEdge (p=0.1) is used to prevent oversmoothing. HyperGCN performs convolution operations (L=3 layers) on the sample-equipment-personnel-reagent hypergraph built during the graph construction phase, learning the graph embedding representations of nodes (especially sample nodes) and modeling the indirect cost propagation path. DropEdge (p=0.1) is used to prevent oversmoothing, and the output graph embedding features are then output.
[0043] In some optional implementations, step S202 analyzes the input vector using a multi-task neural network to obtain a first cost prediction value, which can be achieved through the following steps: Step S2025: Analyze the input vector using the direct cost prediction subnetwork of the multi-task neural network to obtain the direct cost; Step S2026: Analyze the input vector using the indirect cost prediction subnetwork of the multi-task neural network to obtain the indirect cost; Step S2027: Analyze the input vector using the cost contribution prediction subnetwork of the multi-task neural network to obtain an effect prediction value, wherein the effect prediction value represents the magnitude of the effect produced by each biological sample; Step S2028: Concatenate the direct cost, the indirect cost, and the effect prediction value to obtain the first cost prediction value. This method predicts different costs through different branches of the multi-task neural network, thus predicting different cost values based on different features of the input vector, making cost prediction more systematic and efficient.
[0044] In one embodiment of this application, a schematic diagram of the structure of a multi-task neural network is shown below. Figure 4As shown, features are extracted and fused through a dual-branch (temporal encoder + graph convolutional network), and finally, the multi-task head outputs the sample-level direct cost, indirect cost, and QALY prediction value. The multi-task neural network shares FC 1024→512→256; after generating the temporal feature vector, the attribute feature vector, and the graph feature vector in the above steps, feature fusion and multi-task prediction are performed: the temporal feature vector and the graph feature vector are concatenated and fused and deep feature extracted through the shared fully connected layer (FC 1024→512→256). Then, the input is fed into three task-specific towers (sub-networks) in the multi-task neural network: cost_dir: ReLU → FC 64 → Linear output → Single-sample direct cost prediction; cost_ind: Same as above → Predicted indirect costs for a single sample of 67; QALY Tower: Tanh → Linear output → Single-sample QALY contribution prediction (range constrained within [-0.5, 1.5]).
[0045] Final output data: Three predicted values for each sample: cost_dir_pred (direct cost prediction), cost_ind_pred (indirect cost prediction), and QALY_pred (effect prediction).
[0046] The training and tuning of the multi-task neural network are as follows: Stage 1: 310,000 samples from the TCGA public dataset, pre-trained for 30 epochs; Stage 2: fine-tuning of the target structure for 50 epochs, with differential privacy (σ=1.0, C=1.5, δ=10^-5).
[0047] Hyperparameter Bayesian search (Optuna): α:β:γ∈[3:2:1, 6:3:2], final 5:3:2; batch_size ∈ {128, 256, 512}, optimal 256; lr ∈ [1e-4, 5e-4], optimal 2.5e-4 (OneCycle).
[0048] Performance: Training on a single A100 card takes 18 minutes per epoch; validation set MAPE 2.8%, QALY Pearson r=0.91.
[0049] In some optional implementations, step S203 analyzes at least the first cost prediction value using a target residual network to obtain a weight matrix, which can be achieved through the following steps: Step S2031: Obtain a cost driver set, wherein the cost driver set represents the reasons for the cost generated by each of the biological samples; Step S2032: Sample the cost driver set multiple times to obtain multiple sub-driver sets, wherein each sub-driver set is missing one reason, and every two sub-driver sets are different; Step S2033: Analyze the cost driver set and the first cost prediction value using the target residual network to obtain the weight matrix, wherein the target residual network includes a multilayer perceptron neural network and a bidirectional long short-term memory neural network. This method allocates indirect costs through different drivers, achieving a basis-based division of indirect costs, so as to more accurately calculate the cost of each biological sample.
[0050] In one embodiment of this application, this part employs the Shapley value (attribution contribution) approximation method to obtain a predefined set of cost drivers (e.g., usage time of specific equipment, consumption of a batch of reagents, number of operations by a specific person), and calculates the marginal contribution: Monte Carlo sampling, Antithetic Sampling, and Owen nested variance reduction are used to efficiently approximate the Shapley value of each cost driver in the cost driver set. This value measures the average marginal contribution of the driver to the predicted cost (relative to the average cost) when it appears in all possible combinations of drivers. By sampling the cost driver set multiple times (r=576 times), the time complexity is reduced from O(2^n) to O(n log n). Parallel computation is performed on 4 GPUs using PyTorch Lightning and DDP, and processing 1000 samples takes approximately 8 seconds. Finally, the Shapley value (attribution contribution) corresponding to each cost driver is output.
[0051] The target residual network can employ a dual residual network (DRN). The input vector mentioned above can be used as the input to this network. The static feature of the sample is x∈R^d (d=112) (R represents the set of real numbers), and the dynamic operation path feature of the sample is z∈R^(t×k) (t=30 operation steps, k=8-dimensional feature / step) (recording the sequence of equipment, personnel, reagents, environmental conditions, etc. experienced by the sample). Feature extraction steps: Branch 1 (static): Multilayer perceptron (MLP) processes the static feature x, including residual connections. Branch 2 (dynamic): Bidirectional LSTM (Bi-LSTM, hidden=64) processes the dynamic path sequence z, capturing temporal dependencies. Feature fusion and weight generation: The outputs of the two branches are concatenated, passed through a fully connected layer (FC 128), and finally through a Softmax layer to output the amortized weight matrix W. The matrix W_ij represents the proportion of the j-th public cost that the i-th sample should amortize, satisfying ∑W_ij = 1 (for each public cost j). Add `L_explain = λ∥W ⊙Grad_CAM∥1` (λ=0.1) to the training loss function. This forces the generated weight matrix W to be as consistent as possible with the feature importance heatmap extracted from the DRN using Grad-CAM technology, ensuring that the allocation logic can be understood by financial auditors. After training, IoU ≥ 0.85. The final output is the allocated weight matrix W.
[0052] In some optional implementations, step S204 allocates the indirect cost data using the aforementioned weight matrix to obtain a second predicted cost value for each biological sample. This can be achieved through the following steps: Step S2041: Obtain contribution values and contribution directions, where the contribution direction indicates the cause of the cost value, leading to an increase or decrease in the cost value of the biological sample, and the contribution value indicates the degree of increase or decrease. Each contribution value corresponds to a contribution direction. Step S2042: Obtain a preset number of contribution values in descending order to obtain key contribution values, and obtain the contribution direction corresponding to each key contribution value. Step S2043: Divide the indirect costs according to the key contribution values and their corresponding contribution directions to obtain multiple indirect cost subsets, where the contribution directions of the data in each indirect cost subset are consistent. Step S2044: Calculate the product of the weight matrix and each indirect cost subset to obtain the second predicted cost value for each biological sample. This method, through the aforementioned allocation method, can reduce the time spent allocating indirect costs from several hours to seconds, improving cost accounting efficiency.
[0053] In one embodiment of this application, a schematic diagram of the attribution and allocation process for indirect costs is shown below. Figure 5As shown, by utilizing the key cost drivers (contribution values) and their contribution directions identified through Shapley values, and combining them with the allocation weight matrix W generated by DRN, various public expenses (such as equipment depreciation, management salaries, and public area energy consumption) are dynamically and meticulously allocated to each individual beneficiary sample. The direct costs of the sample, the allocated indirect costs, QALY predictions, and the Shapley value attribution explanations and allocation weight basis are integrated to generate a sample-level full lifecycle cost report.
[0054] For each public cost pool (indirect cost), the Shapley values of all drivers are first calculated, and the top-K (usually K=5~7) are selected after sorting by size. These are the "key cost drivers." The "contribution direction" is given by the sign of the Shapley value—a positive value indicates that an increase in the driver will push up public costs, while a negative value indicates that an increase will actually decrease costs, thus clarifying whether this public cost should be allocated "with" or "against" the change in the driver. According to the top-K key drivers given by the Shapley values and their positive and negative contribution directions, the current public cost pool is split into sub-pools with the same direction. The DRN network then takes the static attributes and dynamic operation paths of the samples as input, and outputs the allocation ratio W_ij of each sample in each sub-pool through SoftMax, satisfying ∑W_ij=1 and consistent with the Shapley direction. Multiplying this ratio by the sub-pool amount yields the indirect cost that the sample should bear, and this, along with the direct cost, QALY, and Shapley interpretation, is written into the report, thus completing traceable and refined allocation.
[0055] In some optional embodiments, the method further includes step S206: after calculating the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value, obtaining the current state of the biobank, wherein the biobank includes multiple biological samples, and the current state includes the status of the equipment, biological samples, and personnel of the biobank; step S207: obtaining next cycle plan data, wherein the next cycle plan data represents the processing procedure of biological samples in the next cycle; step S208: constructing a three-dimensional biobank model based on the total cost prediction value, the current state of the biobank, and the next cycle plan data to simulate the real operating environment of the biobank; step S209: performing Monte Carlo budget simulation in the three-dimensional biobank model to obtain budget simulation results; step S210: calculating the error between the total cost prediction value and the budget simulation results, and adjusting the next cycle plan data if the error exceeds a preset threshold to optimize resource scheduling in the next cycle. This method models the real operating environment through the above steps to verify the accuracy of the prediction results.
[0056] In one embodiment of this application, the current state refers to the real-time status data of the current sample library (equipment status, inventory level, personnel location / status). The next cycle plan data refers to the next cycle business plan data (predicted sample volume, testing project plan, personnel shift schedule draft, equipment maintenance plan, etc.). Based on Unity3D + ML-Agents, a 1:1 CAD model (4K PBR texture) at the equipment level is imported to construct a high-fidelity digital twin of the sample library, resulting in a three-dimensional sample library model. The physical world's samples, equipment, personnel, and processes are mapped to the virtual environment. Parameter configuration and perturbation injection are performed on the above model: configuring equipment failure models (Weibull distribution, parameters support PyMC3 online Bayesian updates), personnel efficiency models, reagent batch-to-batch difference models, etc. It supports defining simulation scenarios (such as simulating an 8-hour equipment downtime or a personnel leave). Simulation-driven: the twin runs in 1-minute time steps, simulating the actual operation process of the sample library (sample reception, processing, storage, transportation). Monte Carlo budget simulations are then performed. After each simulation run, the virtual environment generates "actual" cost data for that run. The simulation engine and model within the twin output a simulation sample-level cost report (simulated final settlement) for that scenario, summarizing the total cost / key metrics for that simulation. Error analysis is then performed. If the error Δ > 0.05 (5%), incremental learning of the model is initiated.
[0057] In some optional implementations, step S209, which involves performing Monte Carlo budget simulation in the aforementioned three-dimensional sample library model to obtain budget simulation results, can be achieved through the following steps: Step S2091: Obtain disturbance parameters, and perform multiple samplings on the disturbance parameters to obtain multiple sampling results, wherein the disturbance parameters represent parameters that affect the budget simulation results; Step S2092: Based on each of the above sampling results, perform Monte Carlo budget simulation in the aforementioned three-dimensional sample library model to obtain the budget simulation results corresponding to each of the above sampling results.
[0058] In one embodiment of this application, Monte Carlo budget simulation is performed in a digital twin environment. Latin hypercube sampling (LHS) is used to handle uncertainties and disturbance parameters (such as equipment failure time points, actual reagent efficacy, and sample arrival fluctuations), executing 10^4 independent simulations. After each simulation run, the virtual environment simulates and generates "actual" cost data for that run. The simulation engine and model within the twin output a simulated sample-level cost report (simulated final settlement) for that scenario, summarizing the total cost / key metrics for that simulation.
[0059] In some optional implementations, step S210, which calculates the error between the total cost prediction and the budget simulation result, can be achieved through the following steps: step S2101: generate probability distribution maps corresponding to the total cost prediction and the budget simulation result respectively; step S2102: calculate the distance between the probability distribution map corresponding to the total cost prediction and the probability distribution map corresponding to the budget simulation result, and obtain the error between the total cost prediction and the budget simulation result.
[0060] In one embodiment of this application, the cost results of all 10^4 simulations in the above steps are collected, and a predicted budget distribution π_b for the next cycle cost is generated using kernel density estimation (KDE, Scott bandwidth). When the actual cycle ends, the true settlement distribution π_a is generated. The Cramér-von Mises distance Δ between the two distributions is calculated.
[0061] In some optional implementations, step S210, which adjusts the next cycle plan data, can be achieved through the following steps: step S2103: incrementally learn the three-dimensional sample library model at preset time intervals until the error between the total cost prediction and the budget simulation result is less than the preset threshold.
[0062] In one embodiment of this application, the current environmental state (inventory, personnel shifts, number of samples in transit, remaining turnaround time (TAT) for each task, etc.) is obtained from a digital twin sandbox. The agent outputs a 64-dimensional discrete action, representing a combination of resource scheduling decisions (e.g., assigning personnel A to workstation X, activating backup equipment Y, selecting transportation route Z). A reward (R) mechanism is set: the environment calculates the reward based on the result of the action execution: R = -(cost overrun + 0.8·TAT delay + 0.5·QALY loss). The objective is to maximize the cumulative reward (i.e., minimize cost, delay, and QALY loss). Network: An Actor-Critic structure is used. The Actor network (MLP 1024→512→64) outputs action probabilities, and the Critic network (with the same structure) evaluates the state value. The PPO algorithm is used to update the network parameters. Policy output: The trained PPO policy runs in the sandbox environment, or in a near-real-time scenario, infers and outputs the optimal resource scheduling action A based on the current state S. The final output includes resource scheduling recommendations (specific personnel scheduling adjustments, equipment activation / sharing plans, transportation route selection, etc.).
[0063] A schematic diagram of the model closed-loop learning and training process is shown below. Figure 6As shown, the trigger condition for the incremental learning strategy is the error evaluation result, with an EWC importance threshold of Fisher's diagonal element > 0.001. The preset time interval can be 7 days, with automatic retraining every 7 days and an old task forgetting rate of <2% (evaluation metric: backward transfer). Based on historical budget data and future plans, Monte Carlo simulations are performed in a digital twin sandbox to generate a budget distribution, which is then compared with the actual budget distribution. Reinforcement learning is used to optimize resource scheduling in the simulation. When the budget-budget discrepancy is too large, incremental model learning is triggered, forming a closed loop.
[0064] In some optional implementations, after calculating the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value, the method further includes step S211: obtaining the model gradient of the multi-task neural network, encrypting the model gradient using a public key to obtain an encrypted gradient; step S212: sending the encrypted gradient to a server, whereby the server calculates the sum of the received encrypted gradients to obtain an aggregated encrypted gradient, decrypts the aggregated encrypted gradient using a private key to obtain an aggregated gradient, and adds Gaussian noise to the aggregated gradient to obtain a noisy aggregated gradient, wherein the server is communicatively connected to the biological sample bank; step S213: updating the multi-task neural network using the noisy aggregated gradient. This method enhances data security during transmission through the aforementioned encryption and decryption steps.
[0065] In one embodiment of this application, this portion is a federal compliance audit chain, such as... Figure 7 As shown, the process begins with encrypted aggregation. Each node participating in the federated learning encrypts its local model gradient g_i, generated after training the model using its private data. Specifically, each node encrypts its local gradient g_i using the public key of the Paillier homomorphic encryption algorithm (2048-bit key), resulting in enc(g_i). Before encryption, the gradient is pruned (L2 ≤ 1.5) to meet differential privacy requirements. The encrypted gradient enc(g_i) is transmitted to the federated server via a secure channel. The server uses the homomorphic addition property of Paillier to directly sum the received enc(g_i), obtaining the aggregated encrypted gradient enc(Σg_i). The server decrypts enc(Σg_i) using the Paillier private key, obtaining the aggregated gradient Σg_i. Gaussian noise (σ = 0.75) satisfying (ε,δ)-DP (ε = 1.0, δ = 10^-5) is added to this aggregation result. The global model (including the multi-task neural network and the target residual network) is updated using the aggregated gradient Σg_i + noise with added noise. The final output is the updated global model parameters.
[0066] Next, blockchain notarization is performed, generating a final cost report (JSON format) and its key metadata (model version hash, hyperparameters, training hash, timestamp, institution ID, sample ID range, etc. used when generating the report) by the attribution and allocation module. The cost report JSON data is used to calculate its root hash through a Merkle tree (depth 20) structure. Merkle trees provide efficient data integrity and membership proof. Transaction construction: A transaction is constructed by including the root hash, metadata, and digital signature (institutional private key signature). Consensus and on-chain: The transaction is broadcast and verified in the Hyperledger Fabric consortium blockchain (Raft consensus, 3 Orderer nodes) network. After consensus is reached, it is packaged into a block (approximately 5000 transactions per block, 256KB in size, 0.5-second block interval). The measured TPS is 3200. Final output data: An immutable transaction record written to the blockchain, containing the unique identifier (root hash) of the cost report and associated metadata.
[0067] Cost audit process: Audit query requests (e.g., queries by sample ID) are received via an endpoint. The requester's identity and permissions are verified using OAuth2 + ECDSA P-256 signature verification. Based on the sample ID, the associated cost report is retrieved (potentially from a local database or off-chain storage). The corresponding documented transaction for this report is retrieved from the blockchain (including root hash, metadata, signature, block height, and transaction ID tx_id). The root hash recorded on the chain is verified to match the root hash calculated from the locally retrieved report, ensuring the report has not been tampered with. The entire lifecycle cost chain for this sample is assembled (which can be linked to multiple reports and operational events). The final output is the audit query response, containing the fields: cost_chain (cost detail chain), model_hash, hyperparams, training_hash, signature (report signature), block_height, and tx_id. A complete and verifiable audit trail is provided.
[0068] The environmental configuration for the above-mentioned cost prediction method for biobanks is as follows: Hardware: Local GPU nodes 4×A100 80 GB NVLink; Federation nodes: Pharmaceutical company A (2×V100), Pharmaceutical company B (4×A6000); Network: Dedicated line 10 Gbps + IPSec VPN backup; Software version: Docker 24.0, Kubernetes 1.27, CUDA 12.2.
[0069] Data size: Sample size: 18,462 cases (11,380 tissues, 5,034 blood samples, 2,048 ctDNA samples); Raw data: HL7 FHIR Observation 2.3 GB (before compression); LIMS consumables records 1.1 GB CSV; manual invoices 4.7 GB (TIFF / PDF); storage: TDengine cluster 3×8 TB NVMe SSD, 0.9TB after compression.
[0070] Deployment steps: CI / CD: GitLab Runner + Kaniko to build images, Harbor private repository; HelmChart: one-click deployment of 18 microservices, values.yaml can be configured with resource limits; Federation initialization: Generate Paillier key pairs and write them to the Vault; Consortium blockchain genesis block: Contains 3 hospitals, 2 pharmaceutical companies, and 1 regulatory agency MSP; Gray-scale release: Argo Rollouts is used, with three phases of 5% → 25% → 100%, and a rollback window of 30 seconds.
[0071] Results: Single-sample accounting delay P50 = 47 ms, P99 = 78 ms; 2024Q2 budget variance rate was 4.1%, a decrease of 67% compared to the same period in 2023; Federal Update: New companion diagnostics projects complete model synchronization within 3 days, with cross-institutional MSE < 0.02; The National Health Commission randomly checked 100 samples, achieving a 100% success rate in tracing the source of infection, with no rectification required. GPU utilization: 93% during training, 42% during inference, and during idle periods, the MIG partition is enabled for sharing with other research tasks.
[0072] Risks and Countermeasures: Data drift: Automatic monthly feature distribution monitoring (alarm when Kolmogorov-Smirnov test p<0.05); Privacy leakage: Regular differential privacy budget audits, forced model retirement when ε accumulation >3.
[0073] The embodiments described above in this application use biological sample lifecycle events as the sole time axis. Through sub-minute alignment of multi-source heterogeneous data, hypergraph propagation of indirect costs, and encrypted gradient aggregation, the prediction error of direct / indirect costs for a single sample is ≤3%, and the time for public cost allocation is reduced from several hours to seconds, thereby meeting the needs of real-time, precise, and compliant health economic accounting for high-throughput sample banks.
[0074] It should be noted that the technical solutions in this application are applicable to virtual network environments, and the users described generally refer to "virtual users." Real users can register user accounts on the server through registration to obtain user identities in the network environment. The same user account can log in to the server through different types of client terminals, enabling the server to identify the same user.
[0075] Interactions between the server and the user can be based on user accounts. The data received or sent by the server to the user is also based on the user account; in reality, the user's client, corresponding to the user account, receives or sends data to the server. Furthermore, users can also communicate with each other through their user accounts. Here, "user" can refer to an individual or an organization, such as a company; this application does not impose specific restrictions.
[0076] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0077] It should be noted that some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The operation numbers, such as 201, 202, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should also be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0078] Figure 8 This application provides a schematic diagram of the structure of a cost prediction device for a biobank, as illustrated in an embodiment. Figure 8 As shown, the device may include: The generation unit 801 is used to acquire initial cost data and generate a cost vector based on the initial cost data, wherein the initial cost data represents the costs incurred in processing biological samples; The first analysis unit 802 is used to generate an input vector based on the cost vector, and analyze the input vector through a multi-task neural network to obtain a first cost prediction value. The multi-task neural network is used to analyze at least the temporal and environmental features of the input vector. The environmental features characterize the operating environment of the biological sample. The first cost prediction value includes at least direct costs and indirect costs. The direct costs represent the expenses directly incurred in processing each biological sample, and the indirect costs represent the expenses incurred in processing multiple biological samples. The second analysis unit 803 is used to analyze at least the first cost prediction value through a target residual network to obtain a weight matrix, wherein the target residual network is used to analyze the cause of the cost value of each biological sample, and the weight matrix represents the weight of the indirect cost allocated to each biological sample. The allocation unit 804 is used to allocate the indirect costs through the weight matrix to obtain a second cost prediction value for each biological sample. The calculation unit 805 is used to calculate the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value.
[0079] In some optional implementations, the generation unit includes a parsing module, an image processing module, and a first stitching module. The parsing module is used to acquire structured cost data, parse and store the structured cost data through a protocol adaptation template to obtain first initial cost data, wherein the structured cost data is cost data that follows data storage rules. The image processing module is used to acquire unstructured cost data, perform image processing on the unstructured cost data to convert the unstructured cost data into structured cost data to obtain second initial cost data. The first stitching module is used to stitch the first initial cost data and the second initial cost data to obtain the cost vector.
[0080] In some optional implementations, the first analysis unit includes a first analysis module, a generation module, a second analysis module, and a second splicing module. The first analysis module is used to analyze the cost vector through a time-series analysis network to obtain a time-series feature vector, wherein the time-series feature vector represents the time-series characteristics of the cost vector. The generation module is used to acquire attribute data of the biological sample and generate attribute feature vectors based on the attribute data, wherein the attribute data includes at least sample types, and the attribute feature vectors represent the attribute characteristics of the biological sample. The second analysis module is used to acquire usage information of the biological sample, construct a graph database model based on the usage information, and analyze the usage information of the biological sample based on the graph database model to obtain a graph feature vector, wherein the usage information includes at least the usage environment and users of the biological sample, and the graph feature vectors represent the environmental characteristics of the biological sample. The second splicing module is used to splice the time-series feature vector, the attribute feature vector, and the graph feature vector to obtain the input vector.
[0081] In some optional implementations, the first analysis unit further includes a third analysis module, a fourth analysis module, a fifth analysis module, and a splicing module. The third analysis module is used to analyze the input vector through the direct cost prediction subnetwork of the multi-task neural network to obtain the direct cost; the fourth analysis module is used to analyze the input vector through the indirect cost prediction subnetwork of the multi-task neural network to obtain the indirect cost; the fifth analysis module is used to analyze the input vector through the cost contribution prediction subnetwork of the multi-task neural network to obtain the effect prediction value, wherein the effect prediction value represents the magnitude of the effect produced by each biological sample; the splicing module is used to splice the direct cost, the indirect cost, and the effect prediction value to obtain the first cost prediction value.
[0082] In some optional implementations, the second analysis unit includes a first acquisition module, a sampling module, and a sixth analysis module. The first acquisition module is used to acquire a cost driver set, wherein the cost driver set is a set representing the reasons for the cost generated by each biological sample. The sampling module is used to sample the cost driver set multiple times to obtain multiple sub-driver sets, wherein each sub-driver set is missing one reason, and every two sub-driver sets are different. The sixth analysis module is used to analyze the cost driver set and the first cost prediction value through the target residual network to obtain the weight matrix, wherein the target residual network includes a multilayer perceptron neural network and a bidirectional long short-term memory neural network.
[0083] In some optional implementations, the cost-sharing unit includes a second acquisition module, a third acquisition module, a partitioning module, and a first calculation module. The second acquisition module is used to acquire contribution values and contribution directions, wherein the contribution direction indicates that the cost value is caused by the biological sample, resulting in an increase or decrease in the cost value, and the contribution value indicates the degree of increase or decrease. Each contribution value corresponds to a contribution direction. The third acquisition module is used to acquire a preset number of contribution values in descending order to obtain key contribution values, and to acquire the contribution direction corresponding to each key contribution value. The partitioning module is used to partition the indirect cost according to the key contribution values and the corresponding contribution directions to obtain multiple indirect cost subsets, wherein the contribution directions of the data in the indirect cost subsets are consistent. The first calculation module is used to calculate the product of the weight matrix and each indirect cost subset to obtain the second cost prediction value for each biological sample.
[0084] In some optional embodiments, the apparatus further includes a first acquisition unit, a second acquisition unit, a construction unit, a simulation unit, and an adjustment unit. The first acquisition unit is used to acquire the current state of the biobank after calculating the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value. The biobank includes multiple biological samples, and the current state includes the status of the equipment, biological samples, and personnel in the biobank. The second acquisition unit is used to acquire next cycle plan data, which represents the processing procedure of biological samples in the next cycle. The construction unit is used to construct a three-dimensional biobank model based on the total cost prediction value, the current state of the biobank, and the next cycle plan data to simulate the real operating environment of the biobank. The simulation unit is used to perform Monte Carlo budget simulation in the three-dimensional biobank model to obtain budget simulation results. The adjustment unit is used to calculate the error between the total cost prediction value and the budget simulation results. If the error is greater than a preset threshold, the next cycle plan data is adjusted to optimize resource scheduling in the next cycle.
[0085] In some optional implementations, the simulation unit includes a sampling module and a simulation module. The sampling module is used to acquire disturbance parameters, sample the disturbance parameters multiple times to obtain multiple sampling results, wherein the disturbance parameters represent parameters that affect the budget simulation results. The simulation module is used to perform Monte Carlo budget simulation in the three-dimensional sample library model based on each sampling result to obtain the budget simulation result corresponding to each sampling result.
[0086] In some optional implementations, the calculation unit includes a generation module and a second calculation module. The generation module is used to generate probability distribution maps corresponding to the total cost prediction value and the budget simulation result, respectively. The second calculation module is used to calculate the distance between the probability distribution map corresponding to the total cost prediction value and the probability distribution map corresponding to the budget simulation result, thereby obtaining the error between the total cost prediction value and the budget simulation result.
[0087] In some optional implementations, the adjustment unit includes a learning module for incrementally learning the three-dimensional sample library model at preset time intervals until the error between the total cost prediction and the budget simulation result is less than the preset threshold.
[0088] In some optional embodiments, the apparatus further includes an encryption unit, a decryption unit, and an update unit. The encryption unit is used to obtain the model gradient of the multi-task neural network after calculating the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value, and encrypts the model gradient using a public key to obtain an encrypted gradient. The decryption unit is used to send the encrypted gradient to a server, so that the server calculates the sum of the received encrypted gradients to obtain an aggregated encrypted gradient, decrypts the aggregated encrypted gradient using a private key to obtain an aggregated gradient, and adds Gaussian noise to the aggregated gradient to obtain a noisy aggregated gradient. The server is communicatively connected to the biological sample bank. The update unit is used to update the multi-task neural network using the noisy aggregated gradient.
[0089] Figure 8 The cost prediction device for biobanks can perform Figure 2 The implementation principle and technical effects of the biobank cost prediction method in the illustrated embodiment will not be repeated here. The specific operation methods of each module and unit in the biobank cost prediction device in the above embodiments have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0090] Figure 9 This is a schematic diagram of the structure of one embodiment of a computing device provided in this application. Figure 9 As shown, in practice, the computing device may include a storage component 901 and a processing component 902.
[0091] Storage component 901 is used to store computer programs and can be configured to store various other data to support operation on a computing device. Examples of this data include instructions for any application or method used to operate on the computing device, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0092] Processing component 902, coupled to storage component 901, is used to execute computer programs in storage component 901 for implementing, etc. Figure 2 The method for predicting the cost of a biobank is shown.
[0093] Furthermore, such as Figure 9 As shown, the computing device may also include other components such as a communication component 903, a display component 904, a power supply component 905, and an audio component 906. Figure 9 The diagram only shows some components and does not mean that the device includes only these components. Figure 9 The components shown. Additionally... Figure 9 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the computing device. The computing device in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT (Internet of Things) device, or a server-side device such as a conventional server, cloud server, or server array. If the computing device in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 9 The components within the dashed box; if the computing device in this embodiment is implemented as a conventional server, cloud server, or server array, etc., then it may not include... Figure 9 The component within the dashed box.
[0094] The processing component described above includes one or more processors to execute computer instructions to complete all or part of the steps in the method described above. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the method described above.
[0095] The aforementioned storage components can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0096] The aforementioned communication component is configured to facilitate wired or wireless communication between the device housing the communication component and other devices. The device housing the communication component can access wireless networks based on communication standards, such as mobile communication networks, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0097] The aforementioned display components may include a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0098] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0099] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0100] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.
[0101] Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.
[0102] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0103] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0104] Finally, it should be noted that the above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for predicting the cost of a biobank, characterized in that, include: Obtain initial cost data and generate a cost vector based on the initial cost data, wherein the initial cost data represents the expenses incurred in processing biological samples; An input vector is generated based on the cost vector, and the input vector is analyzed by a multi-task neural network to obtain a first cost prediction value. The multi-task neural network is used to analyze at least the temporal and environmental features of the input vector. The environmental features characterize the operating environment of the biological sample. The first cost prediction value includes at least direct costs and indirect costs. The direct costs represent the expenses directly incurred in processing each biological sample, and the indirect costs represent the expenses incurred in processing multiple biological samples. The first cost prediction value is analyzed at least by a target residual network to obtain a weight matrix, wherein the target residual network is used to analyze the cause of the cost value of each biological sample, and the weight matrix represents the weight of the indirect cost allocated to each biological sample; The indirect costs are allocated using the weight matrix to obtain a second cost prediction value for each biological sample; The total cost prediction value for each biological sample is calculated based on the first cost prediction value and the second cost prediction value.
2. The method according to claim 1, characterized in that, Obtaining initial cost data and generating a cost vector based on the initial cost data includes: Obtain structured cost data, parse and store the structured cost data through a protocol adaptation template to obtain first initial cost data, wherein the structured cost data is cost data that follows data storage rules; Unstructured cost data is acquired, and image processing is performed on the unstructured cost data to convert it into structured cost data, thereby obtaining second initial cost data. The first initial cost data and the second initial cost data are concatenated to obtain the cost vector.
3. The method according to claim 1, characterized in that, Generating an input vector based on the cost vector includes: The cost vector is analyzed by a time series analysis network to obtain a time series feature vector, wherein the time series feature vector represents the time series characteristics of the cost vector; Obtain the attribute data of the biological sample, and generate an attribute feature vector based on the attribute data, wherein the attribute data includes at least the sample type, and the attribute feature vector represents the attribute characteristics of the biological sample; The usage information of the biological sample is obtained, a graph database model is constructed based on the usage information, and the usage information of the biological sample is analyzed based on the graph database model to obtain a graph feature vector. The usage information includes at least the usage environment and the user of the biological sample, and the graph feature vector represents the environmental characteristics of the biological sample. The input vector is obtained by concatenating the temporal feature vector, the attribute feature vector, and the graph feature vector.
4. The method according to claim 1, characterized in that, The input vector is analyzed using a multi-task neural network to obtain a first cost prediction value, including: The direct cost is obtained by analyzing the input vector through the direct cost prediction subnetwork of the multi-task neural network. The indirect cost is obtained by analyzing the input vector through the indirect cost prediction subnetwork of the multi-task neural network. The input vector is analyzed by the cost contribution prediction subnetwork of the multi-task neural network to obtain the effect prediction value, wherein the effect prediction value represents the magnitude of the effect produced by each biological sample; The first cost prediction value is obtained by concatenating the direct cost, the indirect cost, and the predicted effect value.
5. The method according to claim 1, characterized in that, By analyzing at least the first cost prediction value through the target residual network, a weight matrix is obtained, including: Obtain a set of cost drivers, wherein the set of cost drivers is a set representing the reasons for the cost incurred by each of the biological samples; The cost driver set is sampled multiple times to obtain multiple sub-driver sets, wherein each sub-driver set is missing one reason and every two sub-driver sets are different. The target residual network is used to analyze the cost driver set and the first cost prediction value to obtain the weight matrix. The target residual network includes a multilayer perceptron neural network and a bidirectional long short-term memory neural network.
6. The method according to claim 1, characterized in that, The indirect costs are allocated using the weight matrix to obtain a second cost prediction value for each biological sample, including: Obtain contribution value and contribution direction, wherein the contribution direction indicates that the cost value is increased or decreased due to the cause of the cost value, the contribution value indicates the degree of increase or decrease, and each contribution value corresponds to a contribution direction; A preset number of contribution values are obtained in descending order of contribution value to obtain key contribution values, and the contribution direction corresponding to each key contribution value is obtained. The indirect costs are divided according to the key contribution value and the corresponding contribution direction to obtain multiple indirect cost subsets, wherein the contribution direction of the data in the indirect cost subsets is consistent. The second cost prediction value for each biological sample is obtained by multiplying the weight matrix with each of the indirect cost subsets.
7. The method according to claim 1, characterized in that, After calculating the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value, the method further includes: Obtain the current status of the biobank, wherein the biobank includes multiple biological samples, and the current status includes the status of the biobank's equipment, the biological samples, and personnel; Obtain the next cycle plan data, wherein the next cycle plan data is data representing the processing procedure of biological samples in the next cycle; Based on the total cost forecast, the current status of the biobank, and the next cycle plan data, a three-dimensional biobank model is constructed to simulate the real operating environment of the biobank. Monte Carlo budget simulation was performed on the three-dimensional sample library model to obtain the budget simulation results; Calculate the error between the total cost forecast and the budget simulation result. If the error is greater than a preset threshold, adjust the next cycle plan data to optimize resource scheduling for the next cycle.
8. The method according to claim 7, characterized in that, Monte Carlo budget simulation was performed on the three-dimensional sample library model to obtain the budget simulation results, including: Obtain the disturbance parameters, and sample the disturbance parameters multiple times to obtain multiple sampling results, wherein the disturbance parameters represent the parameters that affect the budget simulation results; Based on each of the sampling results, a Monte Carlo budget simulation is performed in the three-dimensional sample library model to obtain the budget simulation result corresponding to each of the sampling results.
9. The method according to claim 7, characterized in that, Calculating the error between the total cost forecast and the budget simulation result includes: Generate probability distribution diagrams corresponding to the total cost prediction and the budget simulation results, respectively; The distance between the probability distribution map corresponding to the total cost prediction and the probability distribution map corresponding to the budget simulation result is calculated to obtain the error between the total cost prediction and the budget simulation result.
10. The method according to claim 7, characterized in that, Adjusting the next cycle plan data includes: Incremental learning is performed on the three-dimensional sample library model at preset time intervals until the error between the total cost prediction and the budget simulation result is less than the preset threshold.
11. The method according to claim 1, characterized in that, After calculating the total cost prediction value for each biological sample based on the first cost prediction value and the second cost prediction value, the method further includes: Obtain the model gradient of the multi-task neural network, and encrypt the model gradient using a public key to obtain the encrypted gradient; The encryption gradient is sent to the server, which calculates the sum of the received encryption gradients to obtain an aggregated encryption gradient. The server then decrypts the aggregated encryption gradient using a private key to obtain an aggregated gradient and adds Gaussian noise to the aggregated gradient to obtain a noise aggregated gradient. The server is connected to the biobank. The multi-task neural network is updated using the noise aggregation gradient.
12. A computing device, characterized in that, This includes processing components and storage components; The storage component stores a computer program; the computer program is invoked and executed by the processing component to implement the cost prediction method for a biobank as described in any one of claims 1 to 11.
13. A computer program product, characterized in that, Includes a computer program or instructions that, when executed by a processing component, implement the cost prediction method for a biobank as described in any one of claims 1 to 11.