Multi-source data fusion and evaluation-based preschool child information ranking system

The system, which integrates multi-source data and intelligent assessment, solves the problems of data dispersion and privacy security in the screening of preschool children for enrollment. It achieves efficient and accurate assessment of enrollment priorities and secure data processing, supporting the scientific allocation of educational resources and the optimization of enrollment policies.

CN121745495APending Publication Date: 2026-03-27FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

The current preschool enrollment survey work suffers from problems such as scattered data sources, inconsistent standards, difficulty in ensuring authenticity, and significant risks to subjective assessment and privacy security. This results in low data integration efficiency and insufficient accuracy, failing to meet the needs of educational resource allocation and enrollment management.

Method used

By constructing a system for trusted fusion and intelligent evaluation of multi-source data, the system uses children's unique identifiers to align entities in data source records, calculates credibility based on the authority of the data source and the evidence chain, conducts evaluation by combining multi-dimensional quantitative features and gradient-enhanced decision tree models, and ensures data security by employing access control and encrypted storage and transmission mechanisms.

Benefits of technology

It significantly improves the efficiency and accuracy of the investigation work, reduces the workload of manual verification, ensures data consistency and privacy security, provides accurate assessment of enrollment priorities, and provides reliable data support for the allocation of educational resources and the optimization of enrollment policies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745495A_ABST
    Figure CN121745495A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source data fusion and evaluation-based preschool child information ranking system, which comprises a data acquisition module, which is configured to acquire ranking data of preschool children in a target area from at least two heterogeneous data sources; the data fusion management module is used for performing entity alignment on records of different data sources based on the unique identifier of the child and detecting conflicts of the same key attribute of the same child; when a conflict exists, calculating the comprehensive credibility of each candidate value based on the credibility weight of a data source for providing a conflict candidate value and the evidence strength of an evidence chain formed by the candidate value and other associated attributes, determining the final fusion value of the key attributes, and generating a standardized module data set; and forming a two-stage collaborative evaluation framework through a rule filtering module and an intelligent sorting module, wherein the rule filtering module executes in advance so as to quickly screen records which do not meet basic conditions, and the intelligent sorting module executes subsequently so as to carry out fine priority evaluation on candidate data which meet the conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer information and data processing technology, specifically relating to a preschool children information survey system based on multi-source data fusion and evaluation. Background Technology

[0002] Currently, the screening process for preschool enrollment still largely relies on the traditional model of manual statistics at the grassroots level, repetitive data collection by multiple departments, and paper-based information verification. The limitations of this model are becoming increasingly apparent in the process of digital transformation in education. From a technical perspective, screening data is often scattered across heterogeneous systems in multiple departments such as education, public security, civil affairs, and health. The lack of unified standards for data formats, field definitions, and coding rules across these systems creates a typical "data silo" phenomenon. This necessitates significant manpower for format conversion and manual cleaning when data is exchanged across departments, resulting in low efficiency and a high risk of mismatches due to semantic ambiguity. Furthermore, the lack of a unified quality verification mechanism for information collected from different channels leads to frequent issues such as duplicate, missing, and abnormal data, directly impacting the completeness and accuracy of the screening data and creating potential problems for subsequent enrollment eligibility verification and resource allocation.

[0003] Even though some regions have made initial attempts at digital methods such as online registration and electronic certificate verification, core technical pain points remain unresolved. On the one hand, the lack of an effective credibility assessment mechanism in the process of multi-source data fusion makes it difficult to objectively determine data priority when key attributes of the same child (such as registered address and length of residence) conflict in different data sources. This still requires reliance on human experience for verification, increasing verification costs and failing to guarantee consistency. On the other hand, existing screening systems focus primarily on information collection, lacking the ability to quantify multi-dimensional features. School enrollment priority judgments often rely on single indicators or subjective experience, failing to comprehensively consider complex factors such as residence consistency and policy compatibility. This results in insufficient accuracy of screening results and cannot provide strong support for the scientific allocation of educational resources.

[0004] Furthermore, data security and privacy protection issues are becoming increasingly prominent. Information about preschool children and their guardians includes sensitive details such as household registration, address, and identification number. Traditional paper storage and manual transmission methods pose a high risk of leakage. Even with digital storage, some systems lack robust access control, data encryption, and operation log traceability mechanisms. This not only fails to comply with relevant personal information protection regulations but may also lead to disputes over educational fairness due to information misuse or illegal acquisition. At the same time, manual screening and verification processes are prone to subjective bias, making it difficult to guarantee the objectivity and fairness of the screening process and failing to meet the current needs of preschool education for precise and equitable enrollment management.

[0005] In summary, traditional screening methods and existing preliminary digital attempts all have significant technical shortcomings in areas such as data fusion standardization, data quality reliability, intelligent assessment and decision-making, and data storage security. Therefore, there is an urgent need to build a technical system that can break down data barriers, ensure data credibility, achieve intelligent assessment, and also protect privacy and security, in order to systematically solve the above problems and promote the development of preschool enrollment screening work towards greater efficiency, accuracy, and standardization. Summary of the Invention

[0006] To address the shortcomings and deficiencies of existing technologies, this invention provides a preschool children's information screening system based on multi-source data credible fusion and intelligent evaluation. This system aims to solve problems such as scattered data sources, inconsistent standards, difficulty in ensuring authenticity, and significant subjective evaluation and privacy security risks in preschool children's enrollment screening. The system integrates heterogeneous data from multiple sources—government channels, online authentication and registration channels, and offline collection channels—through a data acquisition module. It uses a child's unique identifier to align records from different data sources. For conflicts in key attributes of the same child, it dynamically assigns credibility weights based on the authority of the data source, the objectivity of the collection, and the quality of historical data. Combining the strength of the evidence chain consisting of strong and weak corroboration, it calculates the comprehensive credibility of each conflict candidate value to resolve conflicts and generate a standardized screening dataset. The feature engineering module converts the attribute information in the standardized data into multi-dimensional quantitative feature vectors containing temporal satisfaction, spatial consistency, and attribute weights. It can also combine basic features to generate higher-order composite features to capture complex relationships. The system employs a two-stage collaborative evaluation architecture comprised of rule-based filtering and intelligent ranking. First, it quickly filters records that do not meet basic requirements using preset hard criteria. Then, it inputs the feature vectors of the candidate data into a trained gradient boosting decision tree ensemble learning model, outputting objective and accurate admission priority scores. Simultaneously, the system management and security modules ensure data privacy and security through mechanisms such as access control, data anonymization, and encrypted storage and transmission. The results analysis and output module performs multi-dimensional statistical analysis and visualizes student information, synchronizing it with relevant business systems. This invention, through the collaborative design of trusted multi-source data fusion, multi-dimensional quantitative feature construction, and two-stage intelligent evaluation, significantly improves the efficiency, accuracy, and standardization of preschool child screening, providing strong data support for the coordinated allocation of educational resources and the optimization of enrollment policies.

[0007] The specific technical solution adopted by this invention to solve its technical problem is as follows:

[0008] A preschool children information survey system based on multi-source data fusion and evaluation includes:

[0009] The data acquisition module is configured to collect survey data on preschool children in the target area from at least two heterogeneous data sources;

[0010] The data fusion governance module, connected to the data acquisition module, is configured to: perform entity alignment of records from different data sources based on the child's unique identifier, detect conflicts of the same key attribute for the same child; when a conflict exists, calculate the comprehensive credibility of each candidate value based on the credibility weight of the data source providing the conflict candidate value and the corroborating strength of the evidence chain formed by the candidate value and other related attributes, determine the final fusion value of the key attribute, and generate a standardized investigation dataset.

[0011] The feature engineering module, connected to the data fusion and governance module, is configured to convert attribute information in the standardized survey dataset into multi-dimensional quantized feature vectors for subsequent module processing.

[0012] The rule filtering module, connected to the feature engineering module, is configured to: perform Boolean logic judgment on the multidimensional quantized feature vector based on a preset set of hard condition rules, and filter out candidate data that meet the basic conditions;

[0013] The intelligent sorting module, connected to the rule filtering module and the feature engineering module, is configured to: input the multi-dimensional quantized feature vectors corresponding to the candidate data into the trained machine learning model and output the admission priority score of each preschool child;

[0014] The results analysis and output module, connected to the intelligent sorting module, is configured to: perform multi-dimensional statistical analysis based on priority scores, and generate and output a student evaluation report by combining data anonymization and encrypted transmission mechanisms.

[0015] The system management and security module is connected to the above modules and is configured to implement system access control and data security.

[0016] The rule filtering module and the intelligent sorting module constitute a two-stage collaborative evaluation architecture: the rule filtering module executes first to quickly filter records that do not meet the basic conditions, and the intelligent sorting module executes subsequently to perform a fine-grained priority evaluation on the candidate data that meets the conditions.

[0017] Furthermore, the heterogeneous data sources include government channels, online authentication and registration channels, and offline collection channels; the government channels are government platforms, the online authentication and registration channels are the entry points for guardians who have been verified with their real names, and the offline collection channels are on-site collection terminals;

[0018] The data acquisition module supports multiple acquisition methods: automatic synchronization of government platform interface data, online filling and electronic certificate material upload by guardians, on-site entry and certificate verification, batch import of standardized templates, and has basic data cleaning and verification functions.

[0019] Furthermore, the data source credibility weight is dynamically assigned based on the authority of the data source, the objectivity of the collection, and the quality of historical data, and supports periodic fine-tuning based on the historical verification error rate of each data source;

[0020] The overall credibility is calculated as a weighted sum of the data source weight item and the evidence chain supporting item. The data source weight item is determined based on the credibility weight of the data source providing the candidate value, and the evidence chain supporting item is determined based on the evidence strength weight and the degree of matching between the candidate value and the evidence.

[0021] Furthermore, the corroboration strength of the evidence chain is divided into strong corroboration and weak corroboration: strong corroboration is that the conflicting candidate value directly matches the confirmed related attribute value provided by the high-weight data source; weak corroboration is that the conflicting candidate value has a logical relationship or spatial consistency with other attribute values.

[0022] Furthermore, the construction of the multidimensional quantized feature vector includes three basic quantization methods:

[0023] Time-based quantification: Based on the relative relationship between actual satisfaction duration and policy-required duration, the priority of the time dimension is quantified;

[0024] Spatial consistency quantification: Based on the distance between the geographic location corresponding to the attribute and the target location, a distance decay function is used to quantify the spatial matching strength;

[0025] Attribute weight quantification: Based on the classification attribute types of preschool children and their guardians, corresponding priority weights are assigned and quantified.

[0026] Furthermore, the feature engineering module also has a composite feature generation function, which generates higher-order features by combining basic quantization features to capture the complex relationships between different attributes.

[0027] Furthermore, the rule filtering module supports the configuration and maintenance of preset hard conditions, and can mark the specific reasons for failing the filter. The hard conditions include basic requirements related to admission qualifications.

[0028] Furthermore, the machine learning model is a gradient boosting decision tree ensemble learning model; the training dataset of the model is a triplet dataset containing feature vectors, school identifiers and admission labels, and the training adopts a cross-validation strategy that combines stratification by school and time-series partitioning by year.

[0029] Furthermore, the intelligent sorting module supports two evaluation modes: real-time scoring of a single child's record, or batch prediction of the entire survey data.

[0030] Furthermore, the multi-dimensional statistical analysis of the results analysis and output module includes potential student source statistics by preset dimensions such as school, community, and enrollment type, and provides a visualization display function related to student source distribution. The student source evaluation report is synchronized to relevant business systems through an interface.

[0031] Compared with the prior art, the present invention and its preferred embodiments have at least the following beneficial effects:

[0032] This system effectively breaks down barriers between multi-source, heterogeneous data. Through a fusion mechanism combining dynamic credibility weighting and evidence chain corroboration, it systematically solves the problems of scattered data sources, inconsistent standards, and conflicting attributes in existing screening processes. This significantly improves the completeness, consistency, and authenticity of the screening data, reducing the workload and errors of manual verification. By constructing multi-dimensional quantitative features and generating composite features, various qualitative and quantitative conditions related to enrollment are transformed into feature vectors that the model can process. Combined with a two-stage evaluation architecture consisting of rule filtering and intelligent ranking, this replaces the traditional method of screening and prioritization relying on human experience, avoiding subjective bias and improving the objectivity and accuracy of enrollment assessment. The system integrates a full-process data security protection mechanism. Through access control, data anonymization, and encrypted storage and transmission, it effectively reduces the risk of sensitive information leakage, complies with relevant personal information protection regulations, and ensures the security and compliance of the screening process. The overall solution has transformed the preschool children screening process from a traditional manual model to a data-driven model, significantly improving work efficiency and standardization. It provides reliable data support for education departments to coordinate the allocation of educational resources, scientifically optimize enrollment policies, and rationally plan school layouts, thus promoting the development of preschool education enrollment management towards precision and standardization. Attached Figure Description

[0033] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0034] Figure 1 This is a functional diagram of the preschool children's information collection system according to an embodiment of the present invention;

[0035] Figure 2 This is a schematic diagram of the construction process of an embodiment of the present invention. Detailed Implementation

[0036] To make the features and advantages of the present invention more apparent and understandable, specific embodiments are described below in detail:

[0037] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0038] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0039] This invention addresses the technical shortcomings of existing preschool children's information gathering methods, such as multiple and repetitive data collection from various sources, difficulties in verifying authenticity, strong subjectivity in manual screening, and high privacy and security risks. It proposes a solution based on a multi-source data trusted fusion and intelligent evaluation model. By constructing a system including core modules such as data fusion governance, rule filtering, and intelligent sorting, it achieves one-stop collection of preschool children's information, automatic conflict verification, objective quantitative sorting, and secure and controllable management. While ensuring data accuracy and privacy security, it significantly improves the efficiency, objectivity, and decision support capabilities of the gathering process. It is suitable for preschool enrollment gathering in key schools and regional education departments, improving data accuracy and reducing management costs.

[0040] The implementation steps of the preschool children screening scheme provided by this invention are as follows: Figure 2 As shown, it includes:

[0041] Step 1: Define the scope and collect basic data

[0042] The target area, the content to be collected (household registration information, property information, actual residence and residence permit of preschool children and their guardians, etc.), the collection method (interface collection, manual registration, batch import), and the output results (statistical data on potential students for each school, statistical data on potential students for each community, statistical data on potential students for each enrollment type, etc.) should be clearly defined. It should be noted that all data and information proposed in this invention are obtained through legal channels, or the relevant schemes should be carried out under the premise of legal authorization from the relevant authorities, and sufficient confidentiality measures should be provided for the collected data to prevent disclosure to third parties.

[0043] Obtain the previous year's preschool enrollment dataset L and enrollment service conditions S for schools in the target area. i , where i represents the enrollment service condition number (i≤ the total number of enrollment service conditions in the target area).

[0044] The data items in the preschool children's admission dataset L include information such as the preschool children's and their guardians' household registration, property ownership, actual place of residence, enrollment type, and school they attend.

[0045] Admissions Service Conditions S iThe information includes age, date of residency, whether the household registration (residence permit) is within the target area's school district, and whether the property is within the target area's school district.

[0046] Step 2: Construction of a hybrid evaluation model based on multidimensional quantitative features. This step is the offline training phase, in which historical admission data L is used to build and train model Z.

[0047] Multi-dimensional quantitative feature engineering: targeting enrollment service conditions S i Design a set of multidimensional quantization feature mapping functions Φ(L) t ,S i ), will each preschool child L t The degree of matching with each condition is transformed into continuous or discrete numerical features: for example, for the set of enrollment service conditions {S} corresponding to each school in the target area. i}, for each child record L t Calculate its value relative to each condition S i Matching features.

[0048] Time-based satisfaction metric function: For conditions involving time requirements (such as residency duration), calculate the relative satisfaction level. For example, if the policy requires a residency duration of t1, and the actual residency duration for a preschool child is t2, its eigenvalue Φ is defined. time =max(0,(t2-t1) / t1), the larger the value, the earlier the settlement time and the higher the priority;

[0049] Spatial consistency quantification function: For conditions involving geographical location (such as whether the property is within a district), calculate the spatial matching strength. Convert the address to coordinates using geocoding, calculate the distance d to the center of the target school, and define its feature value Φ. location =exp(-λ*d), where λ is the attenuation coefficient greater than 0, and the value is closer to 1 the closer the distance; it can be set according to the radius of the area or experience, for example, it can be set as the reciprocal of the average radius of the area.

[0050] Attribute weight quantization function: For categorical attribute conditions (such as household registration status, home purchase status, etc.), different priority weights p are assigned. For example, "household registration and home purchase" p=1.0, "household registration and home purchase" p=0.8, "household registration and home purchase" p=0.6, and "no home purchase and no household registration" p=0.4. The feature value Φ is defined as follows. property =p;

[0051] Composite feature construction: Combining basic quantized features into higher-order features to capture complex relationships, such as Φ. 两一致 =Φ location (Registered Address) *Φ time (Time of settlement) * I (Registered address = Property address), I(x) means that if condition x is met, it is 1, otherwise it is 0;

[0052] The above feature engineering is applied to the entire preschool children's admission dataset L to obtain a multidimensional feature matrix F, where each preschool child corresponds to a rich feature vector;

[0053] Two-stage hybrid model architecture: Construct a hybrid model Z consisting of a rule filtering layer and an intelligent ranking layer;

[0054] Phase 1: Rule-based filtering layer based on hard conditions, defining a set of hard conditions for preschool children H={h1,h2,...,h...} p}, such as minimum age requirements, document type requirements, etc.; this layer adopts deterministic Boolean logic rules for the preschool children to be assessed. If it violates any of the mandatory conditions If it does not meet the basic requirements, it will be directly marked as "not meeting the basic requirements" and the process will be terminated, and it will not proceed to the next sorting step.

[0055] The second stage: an intelligent ranking layer based on basic learning, which, for preschool children filtered by rules, assigns their multidimensional feature vector F... t Input a gradient boosting decision tree ensemble learning model (such as XGBoost, LightGBM, etc.) to obtain a score for evaluating a student's relative priority within a given school. Finally, the sorted results are output in descending order of scores;

[0056] Model Training and Validation: The feature matrix F is combined with the admission label and school identifier to construct a triplet training dataset D. train =(F(L t ),school(L t ),label(L t ), where label(L t )=1 indicates admission, label(L) t A value of 0 indicates no admission. A cross-validation strategy combining school-based stratification and year-based time-series partitioning is employed to ensure the model's generalization ability across different schools and time series, thereby improving evaluation reliability. As a preferred embodiment, a spatiotemporal cross-validation strategy can be used: first, stratification is performed by school to ensure that samples from each school are distributed in both the training and validation sets; then, in the time dimension, data from earlier years are used as the training set, and data from later years as the validation set, to verify the model's generalization ability over time.

[0057] As a preferred approach, key model parameters are determined through grid search combined with cross-validation, with the learning rate initially set to 0.1 and dynamically adjusted based on validation set performance; tree depth is limited to 3-8 layers to balance model complexity and generalization ability.

[0058] Step 3: Reliability-weighted fusion of multi-source heterogeneous data

[0059] To achieve high-quality integration of the survey data, this step performs credibility-weighted fusion of multi-source heterogeneous data from government platforms (m1), online authentication and registration (m2), and offline collection (m3) to output a standardized survey dataset m with multi-party credibility-weighted fusion.

[0060] Dynamically assigning data source credibility weights: for each original data source m k (k=1,2,3) Assign a basic credibility weight w k The weights are dynamically determined based on the authority of the data source, the objectivity of the collection method, and the quality of historical data (for example, the system can record the frequency with which each data source is adopted as the final value in historical conflict resolution, using this as a feedback indicator of its data quality, and periodically fine-tuning the weight of that data source). Typical assignments are as follows: Government platform data (m1), originating from authoritative departments, has the highest weight, set to w1=1.0; data registered after unified online identity authentication (m2), with traceability, has the next highest weight, set to w2=0.9; data collected offline in communities (m3), involving subjective verification, has a weight set to w3=0.7. Weight w k It supports periodic dynamic fine-tuning of weights based on the historical verification error rate of each source data;

[0061] Conflict detection: First, records from different sources are aligned using a child's unique identifier (such as a hash value of the child's ID card, child's name + guardian's ID number + date of birth) to confirm that they point to the same target. If the same key attribute of the same child (such as "registered address" or "property address") has semantically irreconcilable differences in different data sources, it is considered a conflict. When a unique identifier cannot be directly matched, a fuzzy matching algorithm based on multiple fields such as name, date of birth, and guardian information can be used for auxiliary association, and the matching results are manually reviewed and confirmed.

[0062] Evidence chain construction: When a conflict arises with a certain attribute, the system automatically retrieves records for the child across all other attributes, constructing a logical evidence chain to corroborate each candidate value in the conflict. Evidence chains are divided into two categories:

[0063] Strong corroboration: The conflicting candidate value v directly matches a confirmed associated attribute value provided by another high-weight data source. For example, does the conflicting "property address" completely match the "real estate registration address" of the family member in the government platform data (m1)?

[0064] Weak corroboration: The conflicting candidate value v has a reasonable logical relationship or spatial consistency with other attribute values. For example, is the conflicting "property address" located within the school district of the child's online registered (m2) "intended school", or is it located in the same administrative region as the social security payment location of the child's guardian?

[0065] Evidence-enhanced weighted fusion: For conflicting attributes, the system uses the following quantization algorithm to calculate a comprehensive credibility score for each candidate value v. ,in For data source weight items, This indicates that if the data source m k The value is 1 if a candidate value is provided, and 0 otherwise. As supporting evidence in the chain of evidence, This is the evidence reward coefficient (which can be set to 0.3), used to control the overall impact strength of the evidence chain. This represents the strength weight of the j-th piece of evidence (1.0 for strong evidence and 0.5 for weak evidence). This means that if the candidate value V matches the logical expectation of the j-th piece of evidence, it is 1; otherwise, it is 0.

[0066] After calculating the Score(v) for all candidate values, the value with the highest score is selected as the final merged value for that attribute. If there is a tie for the highest score, the value provided by the data source with the higher weight is adopted first.

[0067] Iterate through all attributes of all surveyed objects, apply the aforementioned conflict detection and fusion decision algorithm, and output a standardized survey dataset m with multi-party trusted weighted fusion.

[0068] Step 4: Model application and result output. This step is the online application stage, where the trained model Z is applied to the standardized data m currently being investigated.

[0069] Using the hybrid assessment model Z trained in step 2, the standardized data m constructed in step 3 is evaluated and analyzed to generate statistical analysis results for the target area by school and by community. These statistical results are then synchronized to the education department or school's business system, including:

[0070] Standardized data input: Based on the multi-dimensional quantitative feature engineering method in step 2, the standardized survey dataset m output by multi-party credible weighted fusion in step 3 is processed to obtain the multi-dimensional feature matrix M of the standardized survey dataset;

[0071] Model evaluation process: For each preschool child L to be evaluated in the survey data m t For children and their intended schools (schools within the designated school district), the rule filtering layer defined in step 2 is invoked to make a judgment based on the hard condition set H; for preschool children who pass the rule filtering, their multidimensional feature vector F is extracted. tInput the pre-trained intelligent ranking layer model to obtain the relative priority score of the preschool child in a given school. ;

[0072] Results generation and output: Aggregate and output all preschool children's assessment results by school and community, and support the synchronization of results to the education department or school business system to support the management and decision-making related to enrollment services.

[0073] Based on the execution and construction process of the above methods, the embodiments of the present invention further design and obtain a preschool children information collection system, such as... Figure 1 As shown, it includes: a registration and screening module, a data fusion and governance module, a feature engineering module, a rule filtering module, an intelligent sorting module, a result analysis and output module, and a system management and security module. Details are as follows:

[0074] Registration and screening module: Serving as the system's data entry point, it enables unified collection and access to multi-source heterogeneous data, including government platform interface acquisition: automatically synchronizing authoritative data such as household registration, residence permits, social security, and real estate from government platforms such as public security, human resources and social security, and natural resources (real estate); Online registration for parents: providing webpage / mini-program entry points for guardians to fill in information, select preferred schools, and upload electronic supporting documents, supporting real-name authentication; Community / kindergarten data collection terminal: providing data entry and verification functions for community staff or kindergarten teachers, supporting on-site information entry, document photography, address location, etc.; Batch import tool: supporting the import of standardized template data in formats such as Excel, and completing basic cleaning and verification.

[0075] Data Fusion Governance Module: Corresponding to "Step 3" in the process, it is responsible for entity alignment, conflict resolution, and credibility fusion of multi-source data, including entity parsing and alignment: based on the child's unique identifier, it associates all records of the same child from different sources; conflict detection and intelligent resolution: when a conflict is detected, it automatically analyzes and determines the final value of key attributes (such as place of household registration, place of property, etc.) based on evidence-enhanced weighted fusion.

[0076] Feature Engineering Module: Corresponding to the multi-dimensional quantization feature engineering in step 2, this module transforms the collected data into numerical feature vectors that the model can understand. It includes a feature calculator with built-in quantization functions (Φ...). time Φ location Φ property Features include: automatic calculation of feature vectors such as time satisfaction and spatial consistency; composite feature generator: automatically combining basic features to generate higher-order features (such as "two-consistency" composite features) according to predefined logic; feature repository: storing all historical feature data for model training and backtracking analysis.

[0077] Rule filtering module: Corresponding to the hard admission criteria in step 2, it provides a fast filtering function, including rule base management: providing a graphical interface to configure and maintain hard condition rules; rule engine: performing rule judgments on each item of the merged data, marking pass, fail and the specific reason for failure.

[0078] Intelligent ranking module: Corresponding to the intelligent ranking layer based on basic learning in step 2, it uses an ensemble learning model to calculate the admission priority score of preschool children filtered by rules. It includes a model service engine: loading and running the trained ranking model (such as XGBoost / LightGBM model), receiving feature vectors, and outputting predicted scores; online / batch prediction: supporting real-time scoring of single records, and also supporting batch prediction of the entire batch of survey data.

[0079] Results Analysis and Output Module: Aggregates and visualizes evaluation results from multiple dimensions and distributes them to relevant business systems. This includes multi-dimensional statistical analysis: automatically generating statistics on the number of potential students by dimensions such as school and community; a visualization dashboard: providing interactive charts to display heat maps of student distribution, trend analysis charts, etc.; and results push and synchronization: pushing statistical results to the business systems of education departments and schools through interfaces.

[0080] System Management and Security Modules: Unified Access Control Module: Implements multi-level role-based access control, including System Monitoring and Auditing Module: Records all data operations and system access logs, and monitors the operating status and performance of each module; Data Security and Privacy Protection Module: Runs through all modules, achieving data anonymization, encrypted transmission and storage, and access logging to ensure compliance with network security and personal information protection regulations.

[0081] The solutions provided in the embodiments of this invention upgrade the preschool child screening process from a traditional manual mode to a data-driven, intelligent, and precise governance system through cross-departmental data fusion and intelligent model evaluation. This effectively solves pain points such as heavy workload, poor data quality, and high security risks. It helps education departments and schools to promptly grasp information such as the number and distribution of students within a target area, providing data support for the overall allocation of educational resources, the scientific optimization of enrollment policies, and the rational planning of school layouts.

[0082] Based on the same inventive concept, this invention also provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the programs include program instructions, and the processor executes the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, used to implement one or more instructions, specifically for loading and executing one or more instructions stored in a computer storage medium to implement the above-described method.

[0083] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, performs the above-described method. This storage medium can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0084] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0085] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

[0086] This invention is not limited to the above-described preferred embodiments. Anyone inspired by this invention can derive other forms of a preschool children information screening system based on multi-source data fusion and evaluation. All equivalent changes and modifications made within the scope of the claims of this invention shall fall within the scope of this invention.

Claims

1. A preschool children's information survey system based on multi-source data fusion and evaluation, characterized in that, include: The data acquisition module is configured to collect survey data on preschool children in the target area from at least two heterogeneous data sources; The data fusion governance module, connected to the data acquisition module, is configured to: perform entity alignment of records from different data sources based on the child's unique identifier, detect conflicts of the same key attribute for the same child; when a conflict exists, calculate the comprehensive credibility of each candidate value based on the credibility weight of the data source providing the conflict candidate value and the corroborating strength of the evidence chain formed by the candidate value and other related attributes, determine the final fusion value of the key attribute, and generate a standardized investigation dataset. The feature engineering module, connected to the data fusion and governance module, is configured to convert attribute information in the standardized survey dataset into multi-dimensional quantized feature vectors for subsequent module processing. The rule filtering module, connected to the feature engineering module, is configured to: perform Boolean logic judgment on the multidimensional quantized feature vector based on a preset set of hard condition rules, and filter out candidate data that meet the basic conditions; The intelligent sorting module, connected to the rule filtering module and the feature engineering module, is configured to: input the multi-dimensional quantized feature vectors corresponding to the candidate data into the trained machine learning model and output the admission priority score of each preschool child; The results analysis and output module, connected to the intelligent sorting module, is configured to: perform multi-dimensional statistical analysis based on priority scores, and generate and output a student evaluation report by combining data anonymization and encrypted transmission mechanisms. The system management and security module is connected to the above modules and is configured to implement system access control and data security. The rule filtering module and the intelligent sorting module constitute a two-stage collaborative evaluation architecture: the rule filtering module executes first to quickly filter records that do not meet the basic conditions, and the intelligent sorting module executes subsequently to perform a fine-grained priority evaluation on the candidate data that meets the conditions.

2. The preschool children information survey system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The heterogeneous data sources include government channels, online authentication and registration channels, and offline collection channels; the government channels are government platforms, the online authentication and registration channels are the entry points for guardians who have been verified with their real names, and the offline collection channels are on-site collection terminals; The data acquisition module supports multiple acquisition methods: automatic synchronization of government platform interface data, online filling and electronic certificate material upload by guardians, on-site entry and certificate verification, batch import of standardized templates, and has basic data cleaning and verification functions.

3. The preschool children information screening system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The data source credibility weight is dynamically assigned based on the data source's authority, collection objectivity, and historical data quality, and supports periodic fine-tuning based on the historical verification error rate of each data source. The overall credibility is calculated as a weighted sum of the data source weight item and the evidence chain supporting item. The data source weight item is determined based on the credibility weight of the data source providing the candidate value, and the evidence chain supporting item is determined based on the evidence strength weight and the degree of matching between the candidate value and the evidence.

4. The preschool children information survey system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The strength of the evidence chain is divided into strong corroboration and weak corroboration: strong corroboration is when the conflicting candidate value directly matches the confirmed related attribute value provided by the high-weight data source; weak corroboration is when the conflicting candidate value has a logical relationship or spatial consistency with other attribute values.

5. The preschool children information survey system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The construction of the multidimensional quantized feature vector includes three basic quantization methods: Time-based quantification: Based on the relative relationship between actual satisfaction duration and policy-required duration, the priority of the time dimension is quantified; Spatial consistency quantification: Based on the distance between the geographic location corresponding to the attribute and the target location, a distance decay function is used to quantify the spatial matching strength; Attribute weight quantification: Based on the classification attribute types of preschool children and their guardians, corresponding priority weights are assigned and quantified.

6. The preschool children information survey system based on multi-source data fusion and evaluation according to claim 5, characterized in that: The feature engineering module also has a composite feature generation function, which generates higher-order features by combining basic quantization features to capture the complex relationships between different attributes.

7. The preschool children information survey system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The rule filtering module supports the configuration and maintenance of preset hard conditions and can mark the specific reasons for failing the filter. The hard conditions include basic requirements related to admission qualifications.

8. The preschool children information survey system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The machine learning model is a gradient boosting decision tree ensemble learning model; the training dataset of the model is a triplet dataset containing feature vectors, school identifiers and admission labels, and the training adopts a cross-validation strategy that combines stratification by school and time-series partitioning by year.

9. A preschool children's information screening system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The intelligent sorting module supports two evaluation modes: real-time scoring of a single child's record, or batch prediction of a whole batch of survey data.

10. A preschool children's information screening system based on multi-source data fusion and evaluation according to claim 1, characterized in that: The multi-dimensional statistical analysis of the results analysis and output module includes potential student source statistics by preset dimensions such as school, community, and enrollment type, and provides a visualization display function related to student source distribution. The student source evaluation report is synchronized to relevant business systems through an interface.