Hospital project configuration type selection scheme automatic generation method based on big data analysis

By constructing an automated big data analysis system, the problems of data silos and lagging resource allocation in modern comprehensive hospital construction projects have been solved. It has realized the intelligent generation of configuration selection schemes from raw data, improved the scientific nature of project decision-making and prediction accuracy, and formed a replicable new model of intelligent construction.

CN121583474APending Publication Date: 2026-02-27YUNJI MEDICAL TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511693808.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Modern comprehensive hospital construction projects suffer from problems such as data silos, lagging resource allocation, uncontrolled progress, cost overruns, and functional redundancy. Existing technologies lack the ability to perceive, intelligently analyze, and collaboratively extrapolate multi-source heterogeneous data in real time, leading to project decisions relying on human experience, severe data silos, a lack of global optimality in solution generation, and a disconnect between resource allocation and clinical practice.

Method used

An automated system based on big data analytics is constructed. Through a unified data lake and standardized preprocessing process, deep integration of multi-source heterogeneous data is achieved. A dynamic weight model that combines the analytic hierarchy process and gradient boosting decision tree is adopted. Multi-objective genetic optimization algorithm and Monte Carlo risk simulation technology are introduced to generate configuration selection schemes that take into account both technical feasibility and economic rationality. A closed-loop feedback mechanism is established to continuously evolve the model.

Benefits of technology

It achieves end-to-end intelligent generation from raw data to complete configuration selection schemes, improves the scientific nature of project decision-making and prediction accuracy, reduces model prediction deviation rate, forms a replicable and scalable new intelligent construction model, supports parallel generation of multiple schemes and interactive decision-making, and improves the transparency of scientific decision-making by management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583474A_ABST
    Figure CN121583474A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, and discloses a hospital project configuration type selection scheme automatic generation method based on big data analysis. The method comprises the steps of multi-source data collection and standardization preprocessing, construction of a dynamic weight evaluation model fusing an analytic hierarchy process and a gradient boosting tree, multi-target Pareto optimization by adopting an improved genetic algorithm, introduction of a Monte Carlo simulation risk pre-judgment and linkage of a resource substitution library adjustment scheme, and output of a structured document and an interactive decision board. The system supports real-time data flow access, elastic space grid planning, post capability matrix matching, dynamic cost modeling and compliance automatic verification. Through a data driving and closed-loop feedback mechanism, automatic generation of a scheme, risk pre-control and continuous evolution of a model are realized, the resource allocation accuracy, the approval passing rate and the implementation efficiency are remarkably improved, and the hyperbranched rate and the modification cost are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence, specifically relating to a method for automatically generating hospital project configuration and selection schemes based on big data analysis. Background Technology

[0002] The research and construction of modern comprehensive hospitals, as the core carriers of the national public health system, ranks among the most complex and system-integrated infrastructure projects in contemporary times. These projects not only involve multiple dimensions such as architectural space planning, medical process reengineering, equipment system integration, and collaborative special projects, but also need to consider clinical service efficiency, scientific research translation capabilities, operational management resilience, and future expansion flexibility. Therefore, they exhibit significant characteristics such as large scale, intensive investment, high degree of professional specialization, and strong dynamic resource allocation. Under traditional project management models, hospital construction often relies on experience-driven linear processes and manual coordination mechanisms. Its core logic lies in achieving project progress through phased approvals and the division of responsibilities. This model could maintain basic operation in the early stages when hospitals were smaller and had relatively simple functions. Its advantages lie in its clear structure, well-defined responsibilities, and predictable implementation path, effectively avoiding the risk of localized loss of control due to information asymmetry.

[0003] However, with the accelerated iteration of medical technology, profound changes in treatment models, and comprehensive upgrades to smart hospital construction standards, the structural defects of the aforementioned traditional models at the principle level are becoming increasingly prominent. The fundamental contradiction lies in the irreconcilable gap between experience-driven static planning mechanisms and data-driven dynamic optimization needs. Specifically, current hospital construction projects generally face practical difficulties such as diverse participating entities, overlapping professional interfaces, a disconnect between equipment procurement cycles and civil engineering progress, and repeated adjustments to functional spaces. The root cause is not simply insufficient management execution, but rather the lack of real-time perception, intelligent analysis, and collaborative deduction capabilities of existing technological systems for massive heterogeneous data. Although hospitals have accumulated a large amount of structured and unstructured data in recent years at various stages of project initiation, budgeting, design, and construction, covering key dimensions such as departmental layout parameters, equipment performance indicators, staffing models, and operating cost curves, this data is mostly scattered in different information systems in silos, lacking a unified semantic framework and dynamic correlation mechanism. When project teams try to extract decision-making basis from the data, they often fall into the paradox of abundant data but scarce information. Manual retrieval and cross-comparison are not only time-consuming and labor-intensive, but also difficult to capture implicit correlation patterns across systems and cycles. This leads to resource allocation plans lagging behind actual needs, which in turn triggers a chain reaction of schedule delays, cost overruns and functional redundancy.

[0004] A deeper technological bottleneck lies in the fact that existing decision support tools mostly remain at the level of data visualization or single-point indicator early warning, failing to build an intelligent inference engine covering the entire chain from demand identification to solution generation, risk prediction, and resource adaptation. For example, in the medical equipment configuration stage, traditional methods usually make linear calculations based on the number of beds or department size, ignoring dynamic variables such as clinical pathway complexity, equipment sharing rate, maintenance cost curves, and technology iteration cycles, resulting in a serious disconnect between the procurement list and the actual usage scenario. At the spatial planning level, static area allocation models are unable to respond to the demand for flexible space in multidisciplinary collaborative diagnosis and treatment models, causing a surge in later renovation costs. In essence, these problems stem from the lack of mathematical modeling capabilities of existing technical architectures for globally optimal solutions under multi-objective constraints—when objective functions such as progress, cost, quality, and compliance are mutually constrained, human experience is difficult to achieve Pareto optimality in high-dimensional parameter spaces, while existing algorithm models, due to insufficient data fusion and crude feature engineering, cannot support accurate decision-making in complex scenarios.

[0005] Therefore, how to construct a system capable of deeply integrating multi-source heterogeneous data, automatically extracting key configuration indicators, dynamically generating configuration selection schemes that balance technical feasibility and economic rationality, and achieving a leap in decision-making paradigm from passive response to proactive prediction has become a core issue in promoting the lean, intelligent, and replicable development of hospital construction projects. This field urgently needs to overcome the dual limitations of existing technologies in terms of the depth of data value mining and the intelligence of scheme generation, and establish a method for automatically generating hospital configuration selection schemes based on big data analysis, driven by artificial intelligence algorithms, and aiming at full lifecycle collaboration, fundamentally reconstructing the decision-making logic and implementation path of hospital construction projects. Summary of the Invention

[0006] This invention aims to address the systemic technical problems in modern comprehensive hospital construction projects caused by the irreconcilable gap between experience-driven static planning mechanisms and data-driven dynamic optimization requirements, leading to issues such as lagging resource allocation, uncontrolled progress, cost overruns, and functional redundancy. Existing technologies lack the ability to perceive, intelligently analyze, and collaboratively extrapolate multi-source heterogeneous data, resulting in project decisions relying on human experience, severe data silos, a lack of global optimality in solution generation, and a disconnect between resource allocation and clinical practice. This invention constructs an automated system covering the entire process of data acquisition, feature extraction, model building, solution extrapolation, and dynamic optimization. It achieves end-to-end intelligent generation from raw data to complete configuration selection solutions, transforming the hospital project configuration selection process from passive response to proactive prediction, from experience-driven to data-model-driven, and from local optimization to global Pareto optimality.

[0007] As one embodiment of the present invention, the method includes the following steps: obtaining raw data from multiple authoritative data sources of target hospitals in China or internationally, including hospital official websites, scientific research project application systems, medical resource management systems, and national and local scientific research project management platforms; importing the obtained raw data into a preset distributed storage system, which adopts a columnar storage structure, supports mixed storage of structured and unstructured data, and establishes a unified data indexing mechanism to ensure that the subsequent data retrieval efficiency is no less than 10,000 query responses per second; performing data cleaning and preprocessing operations on the stored raw data, including removing duplicate records, removing data entries with a field missing rate of more than 30%, performing Z-score standardization on numerical fields, performing one-hot encoding transformation on categorical fields, and performing sliding window smoothing on time series data, ultimately forming a standardized dataset.

[0008] Furthermore, based on a standardized dataset, a core evaluation model for generating hospital configuration selection schemes is constructed. This evaluation model adopts an analytic hierarchy process (AHP) architecture. Its top-level objective is a comprehensive score for hospital project configuration selection schemes. The primary indicators include: rationality of bed allocation, completeness of department setup, staffing matching degree, building area utilization rate, medical equipment adaptability, advancement of medical technology, and standardization of medical management. Each primary indicator has several secondary indicators; for example, medical equipment adaptability includes equipment sharing rate, equipment maintenance cost coefficient, equipment technology iteration cycle, and equipment clinical pathway coverage rate. A judgment matrix is ​​constructed for each level indicator, and expert scoring is performed using a 1-9 scale. The initial weights of each indicator are obtained by calculating the maximum eigenvalue of the judgment matrix and its corresponding eigenvector. A consistency test is performed on the initial weights; if the consistency ratio is less than 0.1, the weight value is retained; otherwise, the judgment matrix is ​​reconstructed until the consistency requirement is met. Finally, the standardized weight vectors of each indicator are output as the basis for calculating subsequent scheme scores.

[0009] Furthermore, a machine learning algorithm is introduced to dynamically optimize the indicator weights. The machine learning algorithm employs a gradient boosting decision tree model, with input features being a 256-dimensional feature vector extracted from a standardized dataset. This feature vector encompasses quantitative indicators such as hospital size parameters, regional economic level, disease distribution density, equipment usage frequency, personnel mobility coefficient, and space turnover rate. Model training utilizes a five-fold cross-validation strategy, with a loss function of Softmax cross-entropy with L2 regularization, a regularization coefficient of 0.01, a learning rate of 0.05, and 500 iterations. After training, the average gain value of each feature in the decision path is extracted and used as the indicator importance score corresponding to that feature. The importance score is then weighted and fused with the initial weights output by the analytic hierarchy process (AHP), with weight fusion coefficients set to 0.6 and 0.4, ultimately outputting a dynamically optimized indicator weight system.

[0010] Furthermore, an automatic generation engine for hospital configuration selection schemes is constructed. The engine receives basic parameters of the hospital to be built from user input, including the planned number of beds, a list of proposed departments, the target service population, budget ceiling, construction period, regional climate conditions, and policy compliance requirements. The engine first calls a dynamically optimized indicator weighting system to perform multi-dimensional mapping of the basic parameters, generating a preliminary configuration framework. Subsequently, the engine initiates a multi-objective optimization module, which employs an improved non-dominated sorting genetic algorithm. Its objective functions include minimizing construction costs, maximizing space utilization, minimizing equipment configuration redundancy, maximizing clinical pathway coverage, and maximizing future expansion flexibility. The algorithm population size is set to 200, the crossover probability to 0.8, the mutation probability to 0.1, and the iteration termination condition is 50 consecutive generations of no change in the optimal solution or reaching the maximum iteration count of 1000. The algorithm outputs a Pareto front solution set, from which the scheme with the highest comprehensive score is selected as the candidate scheme.

[0011] Furthermore, the candidate solutions enter the risk prediction and resource adaptation module. This module constructs a Monte Carlo simulator to simulate 10,000 random disturbances on the key resource configuration items in the candidate solutions. The disturbance range is set based on historical project fluctuation data. For example, the disturbance range for equipment delivery cycle is ±15 days, the disturbance range for personnel attendance rate is ±8%, and the disturbance range for material price fluctuation is ±12%. The simulation outputs the risk probability distribution of each resource configuration item. For items with a risk probability exceeding 20%, an early warning mechanism is triggered, and the resource substitution library is automatically called to adjust the solution. The resource substitution library includes an equipment model substitution matrix, space function replacement rules, and personnel job compatibility mapping table. The substitution adjustment follows the principle of minimizing cost increment. The adjusted solution re-enters the multi-objective optimization module for secondary optimization until the probability of all risk items is lower than the preset threshold.

[0012] Furthermore, the method also includes a solution output and decision support module. This module transforms the final optimized configuration selection scheme into a structured document, which includes a departmental layout diagram, equipment configuration list, staffing table, construction progress Gantt chart, cost breakdown structure table, and risk response plan. The document format supports three standard formats: PDF, DOCX, and XLSX, and embeds digital watermarks and hash check codes to ensure the integrity and traceability of the scheme. At the same time, the module provides an interactive decision dashboard, allowing users to manually adjust any parameter in the scheme. The system responds in real time and recalculates the scheme score and risk value to assist decision-makers in making final confirmations.

[0013] Furthermore, the method establishes a closed-loop feedback mechanism during implementation. During the project implementation phase, actual construction data is collected, including equipment delivery time, construction progress deviations, actual personnel allocation, and detailed cost expenditures. Deviation analysis is performed between the actual data and the predicted data; indicators with deviations exceeding 10% trigger a model retraining process. The retraining process employs an incremental learning strategy, using only new data to fine-tune the original model, avoiding the waste of computational resources caused by full retraining. After the model is updated, it is automatically pushed to all ongoing project systems, achieving knowledge accumulation and continuous evolution of the solution.

[0014] Furthermore, the data acquisition module in the method supports real-time data stream access. By deploying IoT sensing nodes, it collects data on the construction site's environmental temperature and humidity, equipment operating status, personnel location information, and material inventory. The sensing nodes use the LoRa wireless communication protocol, with a communication distance of no less than 3 kilometers and a data upload frequency of once every 5 minutes. After preliminary filtering and aggregation by the edge computing gateway, the data stream is uploaded to the central data lake. The central data lake adopts a stream-batch integrated architecture, supporting joint analysis of real-time and historical data, ensuring that the solution generation engine can dynamically adjust based on the latest site conditions at any time.

[0015] Furthermore, the medical equipment configuration submodule in the method employs a dual-path computation mechanism. The first path is based on clinical pathway analysis, extracting all diagnostic and treatment projects that the target hospital intends to carry out, statistically analyzing the equipment type, usage duration, and concurrent requirements for each project, and calculating the theoretical minimum equipment configuration quantity. The second path is based on big data pattern mining, analyzing equipment configuration data of existing hospitals of similar scale, type, and region, using the K-nearest neighbor algorithm to match similar cases, extracting their equipment configuration mean and standard deviation, and generating a recommended configuration range. The final configuration scheme is the weighted average of the theoretical minimum value and the upper limit of the recommended range. The weight coefficients are dynamically adjusted according to the hospital's financial adequacy; when the financial adequacy is above 80%, the weights are 0.3 and 0.7, and when it is below 50%, the weights are 0.7 and 0.3.

[0016] Furthermore, the spatial planning submodule in the method employs a flexible grid algorithm. The hospital building plan is divided into square grid units with sides of 5 meters. Each grid is assigned a functional attribute label, including treatment area, waiting area, equipment area, passageway area, and buffer zone. The algorithm dynamically adjusts the number of grids and their adjacency relationships in each functional area based on a departmental patient flow density prediction model, ensuring that the peak patient flow density does not exceed 4 people per square meter. At the same time, no less than 15% of the total number of grids is reserved as flexible grids, whose functional attributes can be redefined in the later stages of the project according to actual needs, supporting flexible conversion of spatial functions. The grid adjustment process follows the principle of minimum movement cost, that is, priority is given to adjusting grids in non-load-bearing wall areas to avoid additional costs caused by structural modifications.

[0017] Furthermore, the personnel allocation submodule in the method constructs a job competency matrix. Matrix rows represent job types, including physicians, nurses, technicians, administrators, and logistics personnel; matrix columns represent competency dimensions, including professional qualifications, work experience, language skills, emergency response capabilities, and IT proficiency; matrix elements represent the minimum qualifying value for that job in that competency dimension; the system automatically matches the required job types and quantities based on the departmental setup and service objectives of the hospital to be built, and calls upon the human resources database to screen qualified candidates; if the number of candidates is insufficient, a competency gap warning is activated, and training plans or external recruitment strategies are recommended to ensure strict alignment between personnel allocation and the job competency matrix.

[0018] Furthermore, the cost control submodule in the method employs dynamic cost modeling technology. A cost decomposition structure tree is established, with the root node representing the total cost and child nodes including civil engineering costs, equipment costs, personnel costs, management costs, and contingency funds. Each child node has detailed cost items; for example, equipment costs include purchase price, transportation costs, installation and commissioning costs, training costs, and maintenance deposits. Cost data is sourced from a historical project database and a real-time supplier quotation interface. Based on the current configuration, the system automatically calculates the values ​​of each cost item and adds a risk reserve fund. The reserve fund ratio is dynamically set based on the risk probability output from the Monte Carlo simulation; for every 10% increase in risk probability, the reserve fund ratio increases by 2%. When the total cost exceeds the budget limit, the system automatically initiates a cost optimization program, prioritizing the reduction of non-core equipment configurations and flexible space areas to ensure that core functions are not affected.

[0019] Furthermore, the compliance verification module in the method integrates the latest national and local medical construction standards. The standard provisions are stored in a structured rule format, with each rule including applicable conditions, constraints, compliance thresholds, and penalty clauses. During the scheme generation process, the system performs rule matching on each configuration parameter, such as checking whether the radiology department area is no less than 80 square meters, whether the ICU bed ratio is no less than 5% of the total beds, and whether the waste elevator's load capacity is no less than 1000 kg. If a violation is found, the current scheme generation path is immediately terminated, and an error code and correction suggestions are returned. All compliance verification results are recorded and serve as a mandatory basis for project approval.

[0020] Furthermore, the method supports parallel generation and comparative analysis of multiple schemes. Users can simultaneously input multiple sets of basic parameters, and the system runs multiple scheme generation processes in parallel, outputting a scheme comparison matrix. The matrix dimensions include total cost, construction period, space utilization rate, equipment advancement index, and comprehensive risk score. Users can adjust the weights of each dimension by dragging and dropping, and the system reorders the scheme recommendation list in real time. The final selected scheme can be exported as a complete implementation package, including all technical drawings, procurement lists, construction plans, and acceptance standards, ensuring that the scheme can be directly delivered to the engineering team for execution.

[0021] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0022] 1. This invention completely solves the problems of data silos and information fragmentation in the process of formulating hospital project configuration and selection schemes. By constructing a unified data lake and a standardized preprocessing process, it achieves deep integration and efficient access to multi-source heterogeneous data.

[0023] 2. This invention breaks through the limitations of the traditional experience-driven model and adopts a dynamic weight model that integrates the analytic hierarchy process and gradient boosting decision tree, so that the indicator weights are automatically adjusted according to the data characteristics.

[0024] 3. This invention is the first to introduce multi-objective genetic optimization algorithm and Monte Carlo risk simulation technology into hospital construction projects, so as to achieve global Pareto optimality and risk pre-control in the process of scheme generation.

[0025] 4. The flexible grid space planning mechanism and job competency matrix personnel allocation model constructed by this invention enable the hospital's spatial layout to have dynamic adaptability, avoiding operational efficiency losses caused by improper configuration.

[0026] 5. This invention achieves continuous evolution of the scheme generation model through closed-loop feedback and incremental learning mechanisms. The efficiency of generating new project schemes increases exponentially with the accumulation of historical data, and the model prediction deviation rate decreases year by year. Within three years, it can be stably controlled within 5%, forming a new intelligent construction model that is replicable and scalable.

[0027] 6. The invention supports the parallel generation of multiple solutions and interactive decision dashboards, enabling management to make scientific decisions with the support of quantitative data, and improving the transparency of solution selection. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the overall technical solution architecture proposed in this invention;

[0029] Figure 2 This is a schematic diagram of the core principle framework of the dynamic weight fusion model in this invention;

[0030] Figure 3 This is a logical flowchart of the collaborative deduction of multi-objective genetic optimization and risk prediction in this invention;

[0031] Figure 4 This is a logical flowchart of the collaborative configuration of flexible grid space planning and job competency matrix in this invention;

[0032] Figure 5 This is a schematic diagram of the multi-level interaction relationship and data flow of data acquisition, model training, scheme generation and closed-loop feedback in this invention;

[0033] Figure 6 This is a schematic diagram of the logical framework of the compliance verification and multi-scheme parallel decision support system in this invention. Detailed Implementation

[0034] This embodiment provides a method for automatically generating hospital project configuration and selection schemes based on big data analysis. This method constructs an automated system covering the entire process of data acquisition, feature extraction, model building, scheme deduction, and dynamic optimization, achieving end-to-end intelligent generation from raw data to a complete configuration and selection scheme. The system architecture comprises seven functional layers: a data acquisition layer, a data preprocessing layer, a core evaluation model layer, a scheme generation engine layer, a risk prediction and resource adaptation layer, a scheme output and decision support layer, and a closed-loop feedback layer. Each layer is seamlessly connected and data flows smoothly between them through standardized data interfaces.

[0035] The data acquisition layer executes step S1: acquiring raw data from multiple authoritative data sources from target hospitals both domestically and internationally. Data sources include hospital official websites, research project application systems, medical resource management systems, and national and local research project management platforms. Data acquisition employs a distributed crawler architecture, deploying eight parallel acquisition nodes. Each node is configured with an independent IP proxy pool, with a capacity of at least 500 dynamic IP addresses to ensure the acquisition process is not blocked by the target servers. The acquisition frequency is dynamically adjusted based on the data source update cycle: static data sources are acquired once every 24 hours, and dynamic data sources are acquired once every 4 hours. The total daily data volume is no less than 50GB, and the data formats include JSON, XML, CSV, PDF, DOCX, and XLSX (six types in total). The acquired data is hashed using the SHA-256 hash algorithm to generate unique data fingerprints for subsequent data deduplication and integrity verification.

[0036] The data preprocessing layer executes step S2: importing the acquired raw data into a pre-defined distributed storage system. The storage system uses an Apache Parquet columnar storage structure and is deployed on the Hadoop HDFS distributed file system. The system supports mixed storage of structured and unstructured data; structured data is directly written to Parquet tables, while unstructured data is converted to a structured format after OCR recognition and NLP entity extraction before storage. A unified data indexing mechanism is established, with index fields including four dimensions: data source identifier, collection timestamp, data topic classification, and data fingerprint. The index uses an inverted index structure to ensure that subsequent data retrieval efficiency is no less than 10,000 queries per second. Data cleaning and preprocessing operations are performed on the stored raw data, with the operation process including five sub-steps:

[0037] The first sub-step performs duplicate record removal by performing precise matching based on data fingerprints. If a match is successful, the record with the latest collection timestamp is retained.

[0038] The second sub-step performs missing value processing, removing data entries with a field missing rate exceeding 30%, and filling fields with a missing rate below 30% using multiple imputation, with the imputation model being Bayesian ridge regression.

[0039] The third sub-step performs Z-score standardization on the numeric field, calculated using the following formula: Where μ is the field mean and σ is the field standard deviation;

[0040] The fourth sub-step performs one-hot encoding transformation on categorical fields, with the encoding dimension automatically expanding based on the number of categories, and the maximum encoding dimension not exceeding 128 dimensions;

[0041] The fifth sub-step performs sliding window smoothing on the time series data, with a window size of 7 time units and a Savitzky-Golay filter as the smoothing algorithm, with a polynomial order of 3. After preprocessing, a standardized dataset is formed, containing at least 10 million records and 256 fields.

[0042] The core evaluation model layer executes step S3: It constructs a core evaluation model for generating hospital configuration selection schemes based on standardized datasets. The evaluation model adopts an analytic hierarchy process (AHP) architecture, with the top-level objective being a comprehensive score for the hospital project configuration selection scheme. It includes seven primary indicators: rationality of bed allocation, completeness of department setup, matching degree of staffing, utilization rate of building area, suitability of medical equipment, advancement of medical technology, and standardization of medical management.

[0043] Each primary indicator has 4 to 8 secondary indicators. For example, medical equipment compatibility has 4 secondary indicators: equipment sharing rate, equipment maintenance cost coefficient, equipment technology iteration cycle, and equipment clinical pathway coverage rate. A judgment matrix is ​​constructed for each primary indicator, with a dimension of n×n, where n is the number of indicators at that level. A 1-9 scaling method is used for expert scoring, with 9 experts covering five professional areas: hospital management, medical equipment, architectural engineering, information technology, and financial management. The initial weights of each indicator are obtained by calculating the maximum eigenvalue λ_max of the judgment matrix and its corresponding eigenvector W. The calculation formula is as follows:

[0044]

[0045] Where A is the judgment matrix, w i Let be the i-th component of the eigenvector. Perform a consistency check on the initial weights and calculate the consistency index. The random consistency index RI is obtained by looking up a table based on the matrix dimension; it represents the consistency ratio. If CR is less than 0.1, the weight value is retained; otherwise, the judgment matrix is ​​reconstructed until the consistency requirement is met. Finally, the standardized weight vectors for each indicator are output. The weight vectors have 256 dimensions and correspond one-to-one with the fields of the standardized dataset.

[0046] The core evaluation model layer executes step S4: introducing a machine learning algorithm to dynamically optimize the indicator weights. The machine learning algorithm adopts a gradient boosting decision tree model, and the model framework is built based on the XGBoost library. The input features are 256-dimensional feature vectors extracted from the standardized dataset, covering quantitative indicators such as hospital size parameters, regional economic level, disease distribution density, equipment usage frequency, personnel mobility coefficient, and space turnover rate.

[0047] The model training employs a five-fold cross-validation strategy, with a training set to validation set ratio of 4:1. The loss function is the Softmax cross-entropy function with L2 regularization, where the regularization coefficient is set to 0.01. The learning rate is set to 0.05, the number of iterations is set to 500, the maximum tree depth is set to 6, and the minimum number of samples per leaf node is set to 50. After training, the average gain value of each feature in the decision path is extracted. The gain value is calculated using the following formula:

[0048]

[0049] Where T is the total number of trees, J t Let I be the number of split nodes in the t-th tree. t,j This is an indicator function; it takes a value of 1 when the feature is selected as a splitting feature at node j, and 0 otherwise. t,j The reduction in loss function due to this split. The average gain value is used as the importance score of the indicator corresponding to this feature, and the importance score is normalized to the [0,1] interval. The importance score is weighted and fused with the initial weights output by the analytic hierarchy process, with weight fusion coefficients set to 0.6 and 0.4, and the calculation formula is W. final =0.6×W AHP +0.4×W XGB The final output is a dynamically optimized index weight system, which is automatically updated once a day at 2:00 AM.

[0050] The solution generation engine layer executes step S5: Building an automatic generation engine for hospital configuration and selection solutions. The engine receives basic parameters of the hospital to be built from user input. These parameters include seven core fields: planned number of beds, list of proposed departments, target service population, budget ceiling, construction period, regional climate conditions, and policy compliance requirements.

[0051] The engine first invokes the dynamically optimized metric weighting system to perform multi-dimensional mapping of the basic parameters. The mapping process includes three sub-steps:

[0052] The first sub-step maps the planned number of beds to the bed allocation rationality index. The mapping function is linear interpolation, and the interpolation range is set to [100, 2000] beds based on historical project data.

[0053] The second sub-step maps the proposed department list to the department setup completeness index. The mapping is based on the department association matrix, where the matrix elements represent the synergy coefficient between two departments. A department combination with a synergy coefficient lower than 0.3 triggers an early warning.

[0054] The third sub-step maps the budget ceiling and construction period to cost control and schedule management indicators. The mapping uses a multivariate regression model, and the regression coefficients are dynamically adjusted according to the regional economic level. After the mapping is completed, a preliminary configuration framework is generated, which includes four core parameters: number of departments, total type of equipment, staffing base, and building area base.

[0055] The engine then initiates a multi-objective optimization module, employing an improved non-dominated sorting genetic algorithm implemented within the NSGA-II framework. The objective function comprises five dimensions: minimizing construction costs, maximizing space utilization, minimizing equipment redundancy, maximizing clinical pathway coverage, and maximizing future expansion flexibility. The algorithm population size is set to 200, the crossover probability to 0.8, the mutation probability to 0.1, the crossover operator to simulated binary crossover, the mutation operator to multinomial mutation, and the distribution exponent to 20. The iteration terminates when the optimal solution remains unchanged for 50 consecutive generations or when the maximum number of iterations (1000) is reached. Each generation of the algorithm performs non-dominated sorting and crowding calculation, retaining the Pareto front solution set, with a maximum size of 100. After outputting the Pareto front solution set, the algorithm calculates a comprehensive score for each solution based on a dynamic optimization weight system. The scoring formula is: Score = ∑(w i ×f i ), where w i f is the weight of the i-th objective. i Let be the normalized value for the i-th objective. Select the solution with the highest overall score as the candidate solution; the number of candidate solutions is 1.

[0056] The risk prediction and resource adaptation layer executes step S6: Candidate solutions enter the risk prediction and resource adaptation module. This module constructs a Monte Carlo simulator, implemented using the Python NumPy library. The random number generator employs the Mersenne Twister algorithm, with a fixed seed of 12345 to ensure reproducible results. 10,000 random perturbation simulations are performed on the key resource configuration items in the candidate solutions. These perturbation items include five categories: equipment delivery cycle, personnel attendance rate, material prices, construction progress, and weather impact.

[0057] The disturbance ranges are set based on historical project fluctuation data. Equipment delivery cycle disturbance ranges are ±15 days, following a normal distribution; personnel attendance rate disturbance ranges are ±8%, following a Beta distribution; material price fluctuation ranges are ±12%, following a log-normal distribution; construction progress deviation ranges are ±10%, following a triangular distribution; and weather impact factors range from 0 to 1, following a uniform distribution. The simulation outputs the risk probability distribution for each resource allocation item. Risk is defined as the probability that a resource allocation item deviates from the baseline value by more than a preset threshold. The threshold is set according to the project type; for example, equipment delivery delays exceeding 30 days are considered high risk.

[0058] An early warning mechanism is triggered for items with a risk probability exceeding 20%. Warning levels are divided into three categories: yellow, orange, and red, corresponding to probability ranges of [20%, 40%), [40%, 60%), and [60%, 100%). Upon triggering an early warning, the resource substitution library is automatically invoked to adjust the solution. The resource substitution library contains three data structures: an equipment model substitution matrix (a 200×200 dimensional matrix where element values ​​represent the substitution compatibility between two equipment models; substitution pairs with a compatibility lower than 0.7 are prohibited); a spatial function replacement rule set (a production rule set with 150 rules, each including preconditions and replacement actions); and a personnel job compatibility mapping table (a mapping relationship table between jobs, calculated based on the job capability matrix).

[0059] The substitution adjustment follows the principle of minimizing incremental costs. The incremental cost calculation includes both direct and indirect costs. Direct costs are the price difference in purchasing alternative resources, while indirect costs include training costs and efficiency loss costs. The adjusted solution re-enters the multi-objective optimization module for secondary optimization. The optimization process is repeated until the probability of all risk items is lower than the preset threshold of 15%.

[0060] The solution output and decision support layer executes step S7: The method also includes a solution output and decision support module. This module transforms the final optimized configuration selection scheme into a structured document containing six core categories: Departmental layout diagrams in AutoCAD DWG format, with layers separated into four categories: walls, equipment, pipelines, and annotations; Equipment configuration list in CSV format, with fields including equipment name, model, quantity, unit price, supplier, technical parameters, and maintenance cycle; Staffing table in XLSX format, with worksheets including job title, staffing quantity, competency requirements, salary range, and recruitment channels; Construction progress Gantt chart in Microsoft Project XML format, with no fewer than 500 task nodes and critical paths automatically highlighted in red; Cost breakdown structure table in JSON format, with a tree structure depth of four levels, and leaf nodes representing specific cost items; Risk response plan in PDF format, including five parts: risk description, response measures, responsible person, timeline, and resource requirements. The document supports exporting in three standard formats: PDF, DOCX, and XLSX. During export, a digital watermark and hash checksum are embedded. The digital watermark uses the DCT domain LSB algorithm, and the hash checksum uses the SHA-512 algorithm, with the checksum appended to the end of the document. The module also provides an interactive decision dashboard developed using the React framework. Users can manually adjust any parameter in the plan, including four categories: number of beds, department list, budget ceiling, and construction period. The system responds to adjustment requests in real time, with a response delay of no more than 3 seconds, recalculating the plan score and risk value. The score and risk value change curves are plotted in real time on the dashboard chart area. This assists decision-makers in making final confirmations. Confirmation triggers a plan locking mechanism; once locked, the plan version number automatically increments, and at least 10 historical versions are retained.

[0061] The closed-loop feedback layer executes step S8: The method establishes a closed-loop feedback mechanism during implementation. During the project implementation phase, a data acquisition agent is deployed. This agent is embedded in the project management software and IoT devices, collecting actual construction data including four categories: equipment delivery time, construction progress deviation, actual personnel allocation, and cost expenditure details. Data collection is conducted once daily, with the collection window from 24:00 to 00:30 the following day. Deviation analysis is performed between the actual data and the predicted data, with the deviation calculated as an absolute percentage error. A deviation exceeding 10% triggers a model retraining process. These triggering indicators include four categories: cost deviation, schedule deviation, personnel allocation deviation, and equipment availability deviation. The retraining process employs an incremental learning strategy, using only new data to fine-tune the original model. The amount of fine-tuning data is no less than 5% of the historical data volume of the triggering indicator. During fine-tuning, the underlying model parameters are frozen, and only the top-level classifier weights are updated. The learning rate is set to 1 / 10 of the original training learning rate, i.e., 0.005, and the number of iterations is set to 50. This avoids the waste of computational resources caused by full retraining; full retraining takes approximately 72 hours, while incremental training takes approximately 2 hours.

[0062] After a model update, an update log is automatically generated, containing four fields: update time, update metrics, performance changes, and version number. The update log is pushed to a message queue, which is subscribed to by all projects under construction. Upon receiving an update notification, the system automatically downloads the new model file and restarts the service, enabling knowledge accumulation and continuous solution evolution. The model update frequency is dynamically adjusted based on the number of projects: once a month for a single project, and once a week for more than 10 projects.

[0063] The data acquisition layer executes step S9: The data acquisition module in this method supports real-time data stream access. IoT sensing nodes are deployed, with the sensing node model being a LoRaWAN Class A terminal. Four types of data are collected from the construction site: ambient temperature and humidity, equipment operating status, personnel location information, and material inventory. The ambient temperature and humidity sensor has a range of -40℃ to 85℃ and 0% to 100%RH, with an accuracy of ±0.5℃ and ±3%RH. The equipment operating status sensor collects three parameters: current, voltage, and vibration, with a sampling frequency of 1Hz. Personnel positioning uses UWB technology, with a positioning accuracy of ±10cm. Material inventory is counted using RFID tags, with a tag reading distance of 5 meters.

[0064] The sensing nodes employ the LoRa wireless communication protocol, operating in the 470MHz to 510MHz frequency band with a transmit power of 20dBm and a communication distance of at least 3 kilometers. Data is uploaded every 5 minutes, with each data packet not exceeding 50 bytes. Data flows through an edge computing gateway for initial filtering and aggregation, deployed in the construction site's server room. Filtering rules include outlier removal and data compression; outliers are determined based on the 3σ principle, and data compression uses Delta encoding. The aggregation period is 15 minutes, and the aggregation method is the average value within a time window.

[0065] Data is uploaded to a central data lake, which is built on Apache Kafka and Apache Flink. The Kafka topic has 8 partitions, and the Flink parallelism is set to 16. The central data lake adopts a unified stream and batch processing architecture. The batch processing layer stores historical data, while the stream processing layer processes real-time data. Both layers of data are accessed through a unified SQL interface, ensuring that the solution generation engine can dynamically adjust based on the latest on-site status at any time, with data latency not exceeding 30 seconds.

[0066] The solution generation engine layer executes step S10: The medical device configuration submodule in the method employs a dual-path computation mechanism. The first path is based on clinical pathway analysis, extracting all diagnostic and treatment projects to be carried out by the target hospital from a standardized dataset, with a minimum of 200 projects. Three parameters are calculated for each project: equipment type, usage duration, and concurrency requirements. Usage duration is measured in minutes per instance, and concurrency requirements are measured in the number of devices. The theoretical minimum number of devices to be configured is calculated using the following formula:

[0067]

[0068] Where T total F represents the total project duration. util T is the equipment utilization coefficient. avail For device availability time, U target The target is utilization rate. The second approach, based on big data pattern mining, filters existing hospitals of similar size, type, and location from a historical project database. The filtering criteria are: bed number difference not exceeding 10%, departmental similarity not less than 80%, and regional economic level difference not exceeding one level. The K-nearest neighbor algorithm is used to match similar cases, with K set to 5. Euclidean distance is used as the distance metric, and the feature vector includes four dimensions: hospital size, total equipment value, number of departments, and annual outpatient volume. The mean and standard deviation of equipment configuration are extracted to generate recommended configuration intervals. The lower limit of the interval is the mean minus one standard deviation, and the upper limit is the mean plus one standard deviation. The final configuration scheme is the weighted average of the theoretical minimum and the upper limit of the recommended interval. The weighting coefficient is dynamically adjusted according to the hospital's financial adequacy, calculated using the formula:

[0069]

[0070] When the funding adequacy is above 80%, the weights are 0.3 and 0.7; when it is below 50%, the weights are 0.7 and 0.3; and the weights are linearly interpolated within the 50% to 80% range. The configuration result outputs a device list, which includes five fields: device name, configuration quantity, theoretical basis, recommendation basis, and final weight.

[0071] The scheme generation engine layer executes step S11: The spatial planning submodule in the method employs a flexible grid algorithm. The hospital building plan is divided into square cell grids with sides of 5 meters. The origin of the grid coordinate system is set at the southwest corner of the building, with the X-axis pointing east and the Y-axis pointing north. Each grid is assigned a functional attribute label, with label types including five categories: treatment area, waiting area, equipment area, passageway area, and buffer zone. The labels are stored using integer codes, with values ​​ranging from 1 to 5.

[0072] The algorithm dynamically adjusts the number of grids and their adjacency relationships in each functional area based on a departmental patient density prediction model. The patient density prediction model is an LSTM neural network, with input features including five dimensions: department type, outpatient volume, inpatient volume, surgical volume, and time period coefficient. The output is the hourly patient density value. The adjustment goal is to ensure that the peak-period patient density does not exceed 4 people per square meter, i.e., no more than 100 people per grid during peak periods. The adjustment process uses a simulated annealing algorithm, with an initial temperature set to 1000, a cooling coefficient set to 0.95, and 1000 iterations. Simultaneously, at least 15% of the total grids are reserved as flexible grids. The initial functional attributes of these flexible grids are set as buffers, and their functional attributes can be redefined later in the project based on actual needs. Redefinition is done through the management interface, and the operation records are written to the audit log.

[0073] The system supports flexible conversion of spatial functions, with conversion rules categorized into three types: interchangeable treatment areas and equipment areas, interchangeable waiting areas and buffer zones, and non-convertible passageways. The grid adjustment process follows the principle of minimum movement cost, which includes wall modification costs, pipeline relocation costs, and equipment relocation costs. Cost data is sourced from a historical project database. Priority is given to adjusting grids in non-load-bearing wall areas, which are identified using building structure drawings with an accuracy rate of at least 99%. To avoid additional costs associated with structural modifications, a threshold of 5% of the total construction cost is set.

[0074] The solution generation engine layer executes step S12: the personnel allocation submodule in the method constructs a job competency matrix. The matrix rows represent job types, which include five main categories: physicians, nurses, technicians, administrative staff, and logistics personnel, further subdivided into 32 subcategories. For example, the physician category includes internists, surgeons, and emergency physicians. The matrix columns represent competency dimensions, which include five dimensions: professional qualifications, work experience, language skills, emergency response capabilities, and IT proficiency.

[0075] The matrix elements represent the minimum qualifying score for that position in that competency dimension, expressed as a quantitative score ranging from 0 to 100. Based on the departmental setup and service objectives of the hospital to be built, the system automatically matches the required job types and quantities. The matching algorithm is a rule-based inference engine with 200 rules, each containing three parts: departmental conditions, service volume conditions, and job output. The system also calls upon a human resources database containing 100,000 records, each with four fields: competency score, current status, geographical location, and expected salary.

[0076] The screening criteria are: a competency score no lower than the minimum qualifying value in the matrix, a current hiring status, and a geographical location within 50 kilometers of the target hospital. If the number of candidates is insufficient, a competency gap warning is activated, with a warning threshold of 80% of the job requirements. Training plans or external recruitment strategies are recommended. Training plans include four fields: course name, training duration, training institution, and expected score improvement. External recruitment strategies include three fields: recruitment channel, salary range, and recruitment cycle. Personnel allocation must be strictly aligned with the job competency matrix. Alignment is calculated as the ratio of the average competency score of the actual allocated personnel to the minimum qualifying value in the matrix. A second recruitment process is triggered when the alignment is below 95%.

[0077] The solution generation engine layer executes step S13: The cost control submodule in this method employs dynamic cost modeling technology. A cost decomposition structure tree is established, with the root node representing the total cost. Child nodes include five categories: civil engineering cost, equipment cost, personnel cost, management cost, and contingency fund. Each child node has detailed cost items; for example, equipment cost includes five items: purchase price, transportation cost, installation and commissioning cost, training cost, and maintenance deposit. Cost data comes from a historical project database and a real-time supplier quotation interface. The historical database contains 500 project records, and the real-time interface connects to the APIs of three mainstream building materials and equipment suppliers. Based on the current solution configuration, the system automatically calculates the values ​​of each cost item. The calculation process includes three sub-steps: the first sub-step calculates the basic cost based on the configuration quantity and unit price; the second sub-step calculates the transportation cost based on the transportation distance and weight. The transportation cost calculation formula is:

[0078] Freight=Distance×Weight×Rate;

[0079] Rate is the fee rate coefficient; the third sub-step calculates installation and commissioning fees and training fees based on installation complexity and training time. A risk reserve fund is also added, with the reserve fund ratio dynamically set based on the risk probability output from the Monte Carlo simulation. For every 10% increase in risk probability, the reserve fund ratio increases by 2%, with a base reserve fund ratio of 5%. When the total cost exceeds the budget limit, the system automatically initiates a cost optimization program. The optimization program uses a greedy algorithm, prioritizing the reduction of non-core equipment configurations and flexible space area. The reduction order is based on the cost-benefit ratio, calculated as the ratio of functional importance score to cost. Core functions are ensured to remain unaffected; core functions are defined as those directly related to clinical pathway coverage, and reductions are prohibited when clinical pathway coverage is below 90%.

[0080] The solution generation engine layer executes step S14: The compliance verification module in this method integrates the latest national and local medical construction standards. The standard provisions are stored in structured rule format (JSON), with each rule containing four fields: applicable conditions, constraint objects, compliance threshold, and penalty clauses. There are a total of 300 rules, covering five areas: architectural design, fire protection, environmental protection, medical specialties, and accessibility facilities.

[0081] During the solution generation process, the system performs rule matching on each configuration parameter using the Drools rule engine. For example, it checks whether the radiology department area is no less than 80 square meters, whether the ICU bed ratio is no less than 5% of the total beds, and whether the waste elevator's load capacity is no less than 1000 kg. The matching process is executed in real time, triggering verification immediately after each generated configuration parameter. If a violation is found, the current solution generation path is immediately terminated, and an error code and correction suggestions are returned. The error code is a 6-digit number, with the first 3 digits indicating the standard category and the last 3 digits indicating the specific clause. All compliance verification results are recorded, including five fields: verification time, verification parameters, rule ID, verification result, and processing status. This record serves as a mandatory basis for project approval. The approval system automatically reads the verification records, and solutions that fail verification are prohibited from entering the approval process.

[0082] The solution output and decision support layer executes step S15: the method supports parallel generation and comparative analysis of multiple solutions. Users can input multiple sets of basic parameters simultaneously, with a maximum of 5 sets. The system runs multiple solution generation processes in parallel, each allocated independent computing resources. Resource isolation is achieved through Docker containers, configured with 4 CPU cores and 16GB of memory. The output is a solution comparison matrix, with dimensions including total cost, construction period, space utilization, equipment advancement index, and comprehensive risk score. The matrix is ​​displayed as a heatmap on the decision dashboard, with color gradients from green to red indicating the degree of superiority or inferiority.

[0083] The final selected solution can be exported as a complete implementation package. This package is a ZIP compressed file containing four categories of documents: all technical drawings, procurement lists, construction plans, and acceptance criteria, with a total of no fewer than 100 files. The solution is designed to be directly delivered to the engineering team for execution. The delivery package includes an execution script that automatically verifies file integrity and version consistency; execution will be rejected if verification fails.

[0084] The system and method described in this embodiment, through the detailed technical implementation above, ensure that the generation process of hospital project configuration selection schemes possesses core capabilities such as data-driven, intelligent optimization, risk control, compliance assurance, and continuous evolution. It comprehensively solves systemic technical problems such as resource allocation lag, schedule loss, cost overrun, and functional redundancy in modern comprehensive hospital construction projects, thereby achieving the goal of lean hospital construction.

[0085] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "inclusion" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0086] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for automatically generating hospital project configuration and selection schemes based on big data analysis, characterized in that, include: Raw data is obtained from multiple authoritative data sources from target hospitals both domestically and internationally. These data sources include hospital official websites, research project application systems, medical resource management systems, and national and local research project management platforms. The acquired raw data is imported into a pre-defined distributed storage system. The storage system adopts a columnar storage structure, supports mixed storage of structured and unstructured data, and establishes a unified data indexing mechanism. Perform data cleaning and preprocessing operations on the stored raw data; A core evaluation model for generating hospital configuration selection schemes is constructed based on a standardized dataset. The evaluation model adopts the analytic hierarchy process (AHP) architecture. A machine learning algorithm is introduced to dynamically optimize the indicator weights. The machine learning algorithm adopts a gradient boosting decision tree model. An automatic generation engine for hospital configuration selection schemes is constructed. The engine receives the basic parameters of the hospital to be built input by the user. The engine first calls the dynamically optimized index weight system to perform multi-dimensional mapping of the basic parameters to generate a preliminary configuration framework. Then, a multi-objective optimization module is started. The module adopts an improved non-dominated sorting genetic algorithm. The algorithm outputs a Pareto front solution set, from which the scheme with the highest comprehensive score is selected as the candidate scheme. The candidate solutions are input into the risk prediction and resource adaptation module. This module constructs a Monte Carlo simulator to simulate 10,000 random disturbances on the key resource configuration items in the candidate solutions. The disturbance range is set according to historical project fluctuation data. The simulation outputs the risk probability distribution of each resource configuration item. For items with a risk probability exceeding 20%, an early warning mechanism is triggered and the resource substitution library is automatically called to adjust the solution. The adjusted solution is re-entered into the multi-objective optimization module for secondary optimization until the probability of all risk items is lower than the preset threshold. The solution output and decision support module transforms the final optimized configuration selection solution into a structured document, and the system responds in real time and recalculates the solution score and risk value.

2. The method for automatically generating hospital project configuration and selection schemes based on big data analysis according to claim 1, characterized in that, In the data cleaning and preprocessing operations, fields with a missing rate of less than 30% are filled using multiple imputation, with Bayesian Ridge regression as the imputation model; when performing sliding window smoothing on time series data, the window size is set to 7 time units, the smoothing algorithm uses Savitzky-Golay filter, and the polynomial order is set to 3.

3. The method for automatically generating hospital project configuration and selection schemes based on big data analysis according to claim 1, characterized in that, In the core evaluation model, the expert domains cover five professional areas: hospital management, medical equipment, architectural engineering, information technology, and financial management. The dimensions of the judgment matrix are set according to the number of indicators at that level, and the consistency indicators... The random consistency index RI is obtained by looking up a table based on the matrix dimension; it represents the consistency ratio.

4. The method for automatically generating hospital project configuration and selection schemes based on big data analysis according to claim 1, characterized in that, The gradient boosting decision tree model is built based on the XGBoost library, version 1.7.0; the maximum tree depth is set to 6, and the minimum number of samples in a leaf node is set to 50; the average gain value is calculated using the following formula: Where T is the total number of trees, J t Let I be the number of split nodes in the t-th tree. t,j This is an indicator function; it takes a value of 1 when the feature is selected as a splitting feature at node j, and 0 otherwise. t,j The amount by which the loss function is reduced due to this split.

5. The method for automatically generating hospital project configuration and selection schemes based on big data analysis according to claim 1, characterized in that, In the multi-objective optimization module, the crossover operator uses simulated binary crossover, the mutation operator uses multinomial mutation, and the distribution index is set to 20; each generation performs non-dominated sorting and crowding calculation, retaining the frontier solution set, with a maximum solution set size of 100; the comprehensive score calculation formula is Score=∑(w i ×f i ), where w i f is the weight of the i-th objective. i Let be the normalized value of the i-th objective.

6. The method for automatically generating hospital project configuration and selection schemes based on big data analysis according to claim 1, characterized in that, The Monte Carlo simulator is implemented using the Python NumPy library. The random number generator uses the Mersenne Twister algorithm with a fixed seed of 12345. The equipment delivery cycle disturbance follows a normal distribution, the personnel attendance rate disturbance follows a Beta distribution, the material price fluctuation follows a log-normal distribution, the construction progress deviation follows a triangular distribution, and the weather impact factor follows a uniform distribution. The resource substitution library includes an equipment model substitution matrix, spatial function replacement rules, and personnel job compatibility mapping table. The substitution adjustment follows the principle of minimizing cost increment.

7. The method for automatically generating hospital project configuration and selection schemes based on big data analysis according to claim 1, characterized in that, It also includes establishing a closed-loop feedback mechanism: during the project implementation phase, actual construction data is collected, including equipment delivery time, construction progress deviation, actual personnel allocation, and cost expenditure details; deviation analysis is performed between the actual data and the predicted data, and indicators with a deviation exceeding 10% trigger the model retraining process; the retraining process adopts an incremental learning strategy, using only new data to fine-tune the original model, with the learning rate set to 0.005 and the number of iteration rounds set to 50 rounds; The updated model is automatically pushed to all projects under construction.

8. The method for automatically generating hospital project configuration and selection schemes based on big data analysis according to claim 1, characterized in that, The data acquisition module supports real-time data stream access: by deploying IoT sensing nodes to collect environmental temperature and humidity, equipment operating status, personnel location information, and material inventory at the construction site; the sensing nodes adopt the LoRa wireless communication protocol with a communication distance of no less than 3 kilometers and a data upload frequency of once every 5 minutes; the data stream is initially filtered and aggregated by the edge computing gateway before being uploaded to the central data lake, which adopts a stream-batch integrated architecture with a data latency of no more than 30 seconds.

9. A system for automatically generating hospital project configuration and selection schemes based on big data analysis, characterized in that, include: The raw data acquisition module is used to obtain raw data from multiple authoritative data sources of target hospitals in China or internationally. These data sources include hospital official websites, scientific research project application systems, medical resource management systems, and national and local scientific research project management platforms. A distributed data storage module is used to import the acquired raw data into a preset distributed storage system. The storage system adopts a columnar storage structure, supports mixed storage of structured and unstructured data, and establishes a unified data indexing mechanism. The data preprocessing module is used to perform data cleaning and preprocessing operations on the stored raw data, including removing duplicate records, removing data entries with a field missing rate of more than 30%, performing Z-score standardization on numerical fields, performing one-hot encoding transformation on categorical fields, and performing sliding window smoothing on time series data, ultimately forming a standardized dataset. The core evaluation model construction module is used to build a core evaluation model for hospital configuration selection schemes based on standardized datasets. The evaluation model adopts the analytic hierarchy process (AHP) architecture and outputs standardized weight vectors for each indicator. The dynamic weight optimization module is used to introduce a gradient boosting decision tree model to dynamically optimize the indicator weights and output the dynamically optimized indicator weight system. The automatic solution generation engine module receives the basic parameters of the hospital to be built input by the user, calls the dynamically optimized indicator weight system to generate a preliminary configuration framework, and outputs candidate solutions through the multi-objective optimization module. The risk prediction and resource adaptation module is used to perform Monte Carlo simulation and resource substitution adjustments on candidate solutions, and output optimized solutions with controllable risks. The solution output and decision support module is used to transform the final optimized solution into a structured document and provide an interactive decision dashboard.

10. The automatic generation system for hospital project configuration and selection schemes based on big data analysis according to claim 9, characterized in that, The system also includes a closed-loop feedback module, which is used to collect actual construction data during the project implementation phase, perform deviation analysis and trigger incremental retraining of the model to achieve continuous evolution of the scheme generation model. The system hardware architecture includes a data acquisition terminal, an edge computing gateway, and a central server cluster. The software architecture adopts a microservice design, with inter-service communication using the gRPC protocol. Data storage adopts a hybrid architecture, and the security mechanisms include a network layer firewall, a host layer intrusion detection, an application layer OAuth2.0 authentication and RBAC access control, and a data layer AES-256 encryption and blockchain evidence storage.

Citation Information

Cited By

  • A pre-management method for coal mine safety production

    CN122288408A