Method and system for inferring interpersonal contact patterns within an urban area

By combining small-sample questionnaire surveys and mobile phone location big data with machine learning, the data acquisition and privacy issues of inferring interpersonal contact patterns within cities have been solved, enabling fine-grained contact pattern inference and improving the accuracy of infectious disease models and the effectiveness of public health measures.

CN122245832APending Publication Date: 2026-06-19SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2026-02-04
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies struggle to infer fine-grained interpersonal contact patterns within cities, and existing methods suffer from difficulties in data acquisition, privacy concerns, and insufficient model credibility.

Method used

By designing a small-sample questionnaire survey, we extracted individual demographic and travel activity characteristics. Combined with machine learning models, we used mobile phone location big data to infer individual interpersonal contact patterns under privacy protection. These patterns were then aggregated to the regional level to construct an intra-city interpersonal contact pattern inference system.

Benefits of technology

It enables fine-grained inference of interpersonal contact patterns within cities, improving the accuracy and reliability of inferences, providing more precise data support for infectious disease transmission models, and enhancing the effectiveness of public health interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122245832A_ABST
    Figure CN122245832A_ABST
Patent Text Reader

Abstract

This invention relates to a method for inferring interpersonal contact patterns within urban areas, comprising: Step S1, designing a small-sample questionnaire survey covering demographic and travel activity characteristics, and extracting individual demographic and travel activity characteristics from the survey; Step S2, constructing a machine learning inference model for the number of interpersonal contacts of individuals based on the individual demographic and travel activity characteristics extracted from the small-sample questionnaire survey; Step S3, under the premise of privacy protection, extracting representative individual samples and their individual characteristics from mobile phone location big data for each region, inputting them into the trained inference model, inferring large-scale individual interpersonal contact patterns with regional representativeness, and finally aggregating the individual inference results to the regional level to obtain the interpersonal contact patterns of the entire region. This invention also relates to a system for inferring interpersonal contact patterns within urban areas. This invention enables fine-grained inference from individual characteristics to contact patterns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for inferring interpersonal contact patterns in urban areas. Background Technology

[0002] Interpersonal contact patterns are key parameters in many infectious disease transmission models. Regional interpersonal contact patterns are derived by aggregating representative individual contact patterns within a region, reflecting the average number of contacts between different age groups in different scenarios across the entire region. Individual interpersonal contact patterns, on the other hand, refer to the average number of contacts between an individual and different age groups in different scenarios. Over the past few decades, the main methods for obtaining individual interpersonal contact pattern data in the field of public health have included direct observation, questionnaires, and wearing sensors.

[0003] However, collecting individual contact data based on questionnaires is costly and makes it difficult to collect sufficiently representative samples covering different regions. Therefore, researchers typically build predictive models of individual interpersonal contact patterns based on limited samples. Early research on individual interpersonal contact pattern inference (before 2022) often relied on demographic characteristics such as age, gender, and income. However, individual interpersonal contact patterns are not only influenced by basic, static demographic characteristics but are also closely related to actual travel activity characteristics. Recent research (after 2022) has begun to introduce travel activity characteristics to infer individual interpersonal contact patterns. Some existing methods combine regionally aggregated population flow data (e.g., population inflow / outflow), but they struggle to capture refined individual behavior, limiting model effectiveness. Others rely entirely on extremely high-precision individual movement data (based on GPS positioning, accuracy within 5 meters) for direct contact inference, raising data privacy issues and lacking verification with real contact data, thus casting doubt on the reliability of the results. Furthermore, existing models primarily focus on larger scales, such as national and provincial levels, lacking fine-grained contact pattern inference within cities. Specifically: Traditional methods typically rely on assumed contact patterns, based on the assumption that regions with similar economic and social conditions have similar contact patterns, or employ the assumption of uniform contact. In recent years, scholars have begun to utilize human mobility characteristics to indirectly infer interpersonal contact characteristics. Methodologically, inference methods incorporating mobility characteristics fall into two categories: (1) Methods based on regional mobility characteristics: Most studies are driven by data availability and use regionally aggregated mobility characteristics (such as population inflow / outflow and total commuting volume at the district / county level) as input variables to explain interpersonal contact. The fundamental drawback of this type of method is that interpersonal contact is essentially a micro-behavior of individuals in a specific time and space. Using macro-regional characteristics will inevitably lose key details, resulting in insufficient accuracy and limited explanatory power of the inference results.

[0004] (2) Methods based on individual movement characteristics: A few cutting-edge studies have attempted to use more refined individual movement characteristics (such as GPS trajectory data with meter-level accuracy), which theoretically can better characterize contact potential. However, these methods face significant limitations in practice. On the one hand, refined individual trajectory data is difficult to obtain, involving privacy and cost issues, and lacks universality; on the other hand, the effectiveness of their models generally lacks direct verification with real contact data, leading to doubts about the credibility and uncertainty of the results.

[0005] Furthermore, existing technologies are mostly limited to the national or provincial level, lacking fine-grained inferences about interpersonal contact information within cities. Summary of the Invention

[0006] In view of this, it is necessary to provide a method and system for inferring interpersonal contact patterns in urban areas, which can construct an inference model of the number of interpersonal contacts that integrates demographic and travel activity characteristics, so as to achieve fine-grained regional inference within the city.

[0007] This invention provides a method for inferring interpersonal contact patterns within urban areas. The method includes: Step S1, designing a small-sample questionnaire survey covering demographic and travel activity characteristics, and extracting individual demographic and travel activity characteristics from the survey; Step S2, constructing a machine learning inference model for the number of interpersonal contacts of individuals based on the individual demographic and travel activity characteristics extracted from the small-sample questionnaire survey; Step S3, under the premise of privacy protection, extracting representative individual samples and their individual characteristics from mobile phone location big data, inputting them into the trained inference model, inferring large-scale individual interpersonal contact patterns with regional representativeness, and finally aggregating the individual inference results to the regional level to obtain the interpersonal contact patterns of the entire region.

[0008] Step S1 includes: Step S11: Extract individual demographic characteristics through small-sample questionnaire surveys and mobile phone location big data. Step S12: Based on the individual's travel location on the same day, obtain the characteristics of the individual's travel activities through the survey data.

[0009] Step S2 includes: Step S21: Use machine learning models to infer the mean and median of individual interpersonal contact patterns, and select the specific model with better performance. Step S22: Use the default parameters of each machine learning model as a baseline; then use a random cross-search method to optimize the hyperparameters. Step S23 uses the coefficient of determination, mean absolute percentage error, mean absolute error, and root mean square error as indicators to evaluate the inference performance and error of the model.

[0010] Step S3 includes: Step S31: Extract representative individual samples of different age groups in each region from mobile phone big data to ensure that the representative individual samples can cover people of all ages and match the actual age distribution of the population in the region. Step S32: Input age-representative individual contact data within the region into the trained machine learning model to infer the interpersonal contact patterns of the entire region; based on the output of the machine learning model, generate mean and median statistics reflecting the contact characteristics of the population in the region.

[0011] Step S3 further includes: Step S33: Compare the simulation results under different contact number parameters, analyze the impact of contact heterogeneity on propagation dynamics, and compare them with the traditional method that assumes uniform contact patterns, so as to improve the generalizability and reliability of the propagation model in different regions.

[0012] This invention provides a system for inferring interpersonal contact patterns in urban areas. The system includes an extraction module, a construction module, and an inference module, wherein: The extraction module is used to design a small-sample questionnaire survey covering demographic and travel activity characteristics, and to extract individual demographic and travel activity characteristics from it. The construction module is used to build a machine learning inference model for the number of interpersonal contacts of an individual based on individual demographic characteristics and individual travel activity characteristics extracted from a small sample questionnaire survey. The inference module is used to extract representative individual samples and their individual characteristics from mobile phone location big data under the premise of privacy protection, input them into the trained inference model, infer large-scale individual interpersonal contact patterns with regional representativeness, and finally aggregate the individual inference results to the regional level to obtain the interpersonal contact patterns of the entire region.

[0013] Specifically, the extraction module is used for: Individual demographic characteristics were extracted through small-sample questionnaire surveys and mobile phone location big data. By using survey data, we can obtain the characteristics of an individual's travel activities based on their location on the same day.

[0014] Specifically, the building module is used for: Machine learning models are used to infer the mean and median of individual interpersonal contact patterns, and specific models with better performance are selected from them. The default parameters of each machine learning model were used as a baseline; then, a random cross-search method was used for hyperparameter optimization. The coefficient of determination, mean absolute percentage error, mean absolute error, and root mean square error are used as indicators to evaluate the inference performance and error of the model.

[0015] Specifically, the inference module is used for: Representative individual samples of different age groups in various regions are extracted from mobile phone big data to ensure that the representative individual samples can cover people of all ages and match the actual age distribution of the population in the region. The machine learning model is trained by inputting age-representative individual contact data within the region to infer the interpersonal contact patterns of the entire region. Based on the output of the machine learning model, mean and median statistics reflecting the contact characteristics of the population in the region are generated.

[0016] Specifically, the inference module is also used for: The simulation results under different contact number parameters were compared to analyze the impact of contact heterogeneity on the propagation dynamics. The results were also compared with those of traditional methods that assume uniform contact patterns, in order to improve the generalizability and reliability of the propagation model in different regions.

[0017] This application first extracts individual demographic and travel activity characteristics based on a questionnaire survey of Shenzhen residents' travel activities and interpersonal contacts. Then, it applies machine learning algorithms to construct an interpersonal contact number inference model that integrates demographic and travel activity characteristics, thereby achieving fine-grained regional inference within the city. In other words: This application is able to extract consistent individual demographic and travel activity characteristics from small-sample questionnaire survey data and mobile phone location big data. Then, using these characteristics, machine learning algorithms are applied to construct an inference model for the number of interpersonal contacts an individual has, achieving accurate inference from individual characteristics to contact patterns. Attached Figure Description

[0018] Figure 1 This is a flowchart of the method for inferring interpersonal contact patterns in urban areas according to the present invention; Figure 2 This is a schematic diagram of the framework of an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the verification of the SEIR infectious disease transmission model according to an embodiment of the present invention; Figure 4 This is a hardware architecture diagram of the urban internal area interpersonal contact pattern inference system of the present invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0020] See Figure 1The diagram shown is a flowchart of a preferred embodiment of the method for inferring interpersonal contact patterns in urban areas according to the present invention. Please refer to it as well. Figure 2 .

[0021] Step S1: Extract individual characteristics. This involves designing a small-sample questionnaire covering demographic and travel activity characteristics, and extracting individual demographic and travel activity characteristics from it. These individual characteristics include both demographic and travel activity features.

[0022] In this embodiment, with the assistance of the Shenzhen Center for Disease Control and Prevention and community staff, a questionnaire survey was conducted at various community workstations. The questionnaire covered residents' basic statistical information, travel frequency and locations, and interpersonal contact information such as the number of people and places they came into contact with. A total of 3022 valid questionnaires were collected, based on which individual demographic characteristics and individual travel activity characteristics were extracted. Specifically: Step S11, extract individual demographic features: In the process of extracting individual demographic features, ensuring the consistency of features from different data sources is crucial for model input. Information such as gender, age, place of residence, and school / work location is directly obtained through small-sample questionnaires and mobile phone location big data. The questionnaire survey revealed that there are no significant differences in contact patterns between genders, and their contribution to model differentiation is limited. To optimize model complexity and generalization ability, gender variables are not included in feature construction in this embodiment. Since income information is lacking in mobile phone location data, this embodiment uses an individual affluence index label as a substitute to indirectly infer their economic level.

[0023] Table 1 Individual Demographic Characteristics

[0024] Date type*: Weekday / Weekend.

[0025] Note: √ indicates that features can be extracted directly from the original dataset.

[0026]

[0027] Step S12, extract individual travel activity features: By analyzing survey data, we obtained residents' travel activity characteristics, such as travel scenarios, duration, frequency, and distance, providing a direct description of individual travel behavior patterns. Simultaneously, we utilized mobile phone location data to further extract residents' actual travel activity characteristics. To ensure the consistency of features extracted from the two data sources and for use in model inference, we calculated individual travel activity characteristics based on the individual's daily travel locations, reflecting the individual's travel frequency, activity range, and movement patterns. This provides data support for further modeling. The following are the definitions of individual travel activity characteristics: Radius of gyration: Radius of gyration is commonly used to measure the range of an individual's activity. The formula for calculation is:

[0028] in: Radius of gyration; The total number of trips; For the first The displacement distance of a trip (usually referring to the straight-line distance from the starting point to the destination).

[0029] Farthest distance from home: The maximum geographical distance between an individual's home address and all travel destinations (or trajectory points) within a day, calculated using the following formula:

[0030] in: It represents the total number of trips an individual makes within a specified time period.

[0031] Mobility entropy: Mobility entropy is an indicator that measures the complexity and diversity of an individual's travel behavior. The formula for its calculation is:

[0032] in: This is the moving entropy.

[0033] The number of travel destinations.

[0034] For individual access Duration of time to each destination.

[0035] Daily activity frequency: The total number of all types of activities that an individual engages in on a given day.

[0036] Attendance Status Indicator for the Day: Used to indicate whether an individual attended school or work on the statistical day. (Value: 0 - No, 1 - Yes) Daily commute distance: The distance an individual travels to or from school on the day the statistics are compiled.

[0037] Frequency of other activities on the same day: The total number of other activities an individual engages in on the day of the statistics, excluding going to school or work.

[0038] Other Activities Status Indicator for the Day: This indicates whether an individual had any activities other than going to school or work on the day of the survey. (Value: 0 - No, 1 - Yes) Table 2 Characteristics of Individual Travel Activities

[0039] Note: √ indicates that features can be extracted directly from the original dataset.

[0040] Step S2: Construct an individual interpersonal contact number inference model: Based on the individual demographic characteristics and individual travel activity characteristics extracted from a small sample questionnaire survey, construct a machine learning inference model for the number of individual interpersonal contacts.

[0041] In this embodiment, each community resident extracted from a small-sample questionnaire survey is considered a sample, and machine learning methods are used to solve the problem of inferring individual interpersonal contact patterns. First, the samples are divided into training and testing sets, and individual demographic characteristics and individual travel activity characteristics are used as input variables to train the machine learning model. Based on this, an inference model for individual interpersonal contact patterns is constructed to infer the number of contacts an individual has with people of different age groups in different scenarios. Through this inference model, the individual's contact patterns are accurately inferred, providing important data support for subsequent research on regional interpersonal contact patterns. Specifically: Step S21, Model Definition: This embodiment uses a machine learning model to infer the mean and median of individual interpersonal contact patterns, and selects the specific model with better performance. The specific definition of the model is as follows: Random Forest Model: Random forest is an ensemble learning method that enhances the inference accuracy of a model by constructing multiple decision trees and combining their inference results. Each tree is trained based on a random subset of samples and a random subset of features. Random forests have strong resistance to overfitting and achieve good results in both classification and regression tasks.

[0042] XGBoost (eXtreme Gradient Boosting) model: XGBoost is a scalable gradient boosting decision tree algorithm. It effectively prevents overfitting by introducing a regularization term into the objective function to control model complexity. XGBoost performs a second-order Taylor expansion of the objective function and supports feature parallelism and data parallelism, resulting in excellent performance in both accuracy and computational efficiency.

[0043] LightGBM (Light Gradient Boosting Machine) model: LightGBM is an efficient gradient boosting framework. It employs a histogram-based decision tree algorithm and introduces two techniques: gradient one-sided sampling (GOSS) and mutually exclusive feature binding (EFB). This significantly improves training speed and reduces memory consumption while maintaining accuracy, making it particularly suitable for large-scale datasets.

[0044] CatBoost (Categorical Boosting) model: CatBoost is a machine learning algorithm based on gradient boosting decision trees. Its core advantage lies in its ability to efficiently and directly process categorical features without extensive preprocessing. By using key techniques such as ordered boosting and target statistics, it effectively avoids gradient bias and prediction offset, demonstrating excellent performance on datasets containing a large number of categorical features.

[0045] In other embodiments, the machine learning algorithm can be changed from the random forest algorithm to other machine learning or deep learning algorithms, such as support vector regression, feedforward neural networks, and elastic networks.

[0046] Table 3 Model List

[0047] Step S22, Parameter Setting and Optimization: This embodiment performs model parameter tuning in two stages: First, the default parameters of each machine learning model are used as a baseline; then, a random cross-reference method is used to optimize hyperparameters to improve model performance. Different loss functions are tested on the mean prediction model to better infer the right-skewed data distribution.

[0048] Table 4. List of Loss Functions for Mean Prediction Model

[0049] Randomized SearchCV (Randomized Crossover Search): Randomized crossover search is an efficient hyperparameter optimization algorithm. Unlike grid search, which traverses all possible combinations, Randomized crossover search randomly selects and evaluates a fixed number of parameter combinations from a specified parameter distribution. It can discover high-performance parameter combinations with a higher probability and in a shorter time, making it particularly suitable for high-dimensional parameter spaces. Its basic process is as follows: Parameter space: Specifies a probability distribution (such as uniform distribution, log-uniform distribution, etc.) for each hyperparameter, rather than a discrete list of values.

[0050] Random sampling: randomly selecting from this parameter space Group hyperparameter configuration, among which The pre-set number of iterations (calculation budget).

[0051] Evaluation and selection: Evaluate the model performance corresponding to each set of parameter configurations on the validation set, and finally select the hyperparameter combination with the best performance.

[0052] This process can be formally represented as finding the optimal hyperparameter configuration. :

[0053] in: This represents a set of hyperparameter configurations.

[0054] It is a predefined joint probability distribution of hyperparameters.

[0055] Is it using hyperparameter configuration? The performance evaluation function of the model on the validation set (such as the coefficient of determination) ).

[0056] It represents the total number of random samples (for budget calculation).

[0057] This indicates the search for a performance evaluation function. Maximize hyperparameter configuration .

[0058] Step S23, Evaluation Indicators: This embodiment uses the coefficient of determination (R²), mean absolute percentage error (MAPE), mean absolute error (MAE), and root mean square error (RMSE) as indicators to evaluate the inference performance and error of the model, and their specific definitions are as follows: Table 5. List of Model Evaluation Indicators

[0059] Step S3: Inferring regional interpersonal contact patterns based on big data. That is, under the premise of privacy protection, representative individual samples and their individual characteristics are extracted from mobile phone location big data for each region, input into a trained model, and inferred large-scale individual interpersonal contact patterns representative of the region. Finally, the individual inference results are aggregated to the regional level to obtain the interpersonal contact patterns of the entire region. Specifically: Step S31, Representative Sample Extraction: This embodiment uses Shenzhen as an example to illustrate the process. Representative individual samples of different age groups in various regions are extracted from mobile phone big data to ensure that these samples cover all age groups and match the actual age distribution of the regional population. For infants and toddlers aged 0-6, whose mobile phone usage is relatively low, individual samples from small-sample questionnaires are used to supplement the data and compensate for any missing information. Ultimately, 1.66 million mobile phone users were extracted.

[0060] Step S32, Inference of regional interpersonal contact patterns: This embodiment inputs age-representative individual contact data within a region into a trained machine learning model to infer interpersonal contact patterns across the entire region. Based on the output of the machine learning model, mean and median statistics reflecting the contact characteristics of the population in the region are generated, providing a crucial data foundation for the analysis and modeling of infectious disease transmission dynamics.

[0061] Step S33, Assessment of Regional Interpersonal Contact Patterns: Please also refer to Figure 3 This embodiment constructs an SEIR transmission model for typical respiratory infectious diseases (such as influenza and COVID-19). By comparing the simulation results under different contact number parameters, it analyzes the impact of contact heterogeneity on transmission dynamics and compares it with the traditional method that assumes uniform contact patterns, thus verifying the advantages of this application and improving the generalizability and reliability of the transmission model in different regions.

[0062] See Figure 4 The diagram shown is a hardware architecture diagram of the interpersonal contact pattern inference system 10 for urban areas according to the present invention. The system includes: an extraction module 101, a construction module 102, and an inference module 103.

[0063] The extraction module 101 is used to extract individual characteristics. That is, to design a small-sample questionnaire survey covering demographic and travel activity characteristics, and to extract individual demographic and travel activity characteristics from it.

[0064] In this embodiment, with the assistance of the Shenzhen Center for Disease Control and Prevention and community staff, a questionnaire survey was conducted at various community workstations. The questionnaire covered residents' basic statistical information, travel frequency and locations, and interpersonal contact information such as the number of people and places they came into contact with. A total of 3022 valid questionnaires were collected, based on which individual demographic characteristics and individual travel activity characteristics were extracted. Specifically: The extraction module 101 extracts individual demographic features: In the process of extracting individual demographic features, ensuring the consistency of features from different data sources is crucial for model input. Information such as gender, age, place of residence, and school / work location is directly obtained through small-sample questionnaires and mobile phone location big data. The questionnaire survey revealed that there are no significant differences in contact patterns between genders, and their contribution to model differentiation is limited. To optimize model complexity and generalization ability, gender variables are not included in feature construction in this embodiment. Since income information is lacking in mobile phone location data, this embodiment uses an individual affluence index label as a substitute to indirectly infer their economic level.

[0065] Table 1 Individual Demographic Characteristics

[0066] Date type*: Weekday / Weekend.

[0067] Note: √ indicates that features can be extracted directly from the original dataset.

[0068]

[0069] The extraction module 101 extracts individual travel activity characteristics: By analyzing survey data, we obtained residents' travel activity characteristics, such as travel scenarios, duration, frequency, and distance, providing a direct description of individual travel behavior patterns. Simultaneously, we utilized mobile phone location data to further extract residents' actual travel activity characteristics. To ensure the consistency of features extracted from the two data sources and for use in model inference, we calculated individual travel activity characteristics based on the individual's daily travel locations, reflecting the individual's travel frequency, activity range, and movement patterns. This provides data support for further modeling. The following are the definitions of individual travel activity characteristics: Radius of gyration: Radius of gyration is commonly used to measure the range of an individual's activity. The formula for calculation is:

[0070] in: Radius of gyration; The total number of trips; For the first The displacement distance of a trip (usually referring to the straight-line distance from the starting point to the destination).

[0071] Farthest distance from home: The maximum geographical distance between an individual's home address and all travel destinations (or trajectory points) within a day, calculated using the following formula:

[0072] in: It represents the total number of trips an individual makes within a specified time period.

[0073] Mobility entropy: Mobility entropy is an indicator that measures the complexity and diversity of an individual's travel behavior. The formula for its calculation is:

[0074] in: This is the moving entropy.

[0075] The number of travel destinations.

[0076] For individual access Duration of time to each destination.

[0077] Daily activity frequency: The total number of all types of activities that an individual engages in on a given day.

[0078] Attendance Status Indicator for the Day: Used to indicate whether an individual attended school or work on the statistical day. (Value: 0 - No, 1 - Yes) Daily commute distance: The distance an individual travels to or from school on the day the statistics are compiled.

[0079] Frequency of other activities on the same day: The total number of other activities an individual engages in on the day of the statistics, excluding going to school or work.

[0080] Other Activities Status Indicator for the Day: This indicates whether an individual had any activities other than going to school or work on the day of the survey. (Value: 0 - No, 1 - Yes) Table 2 Characteristics of Individual Travel Activities

[0081] Note: √ indicates that features can be extracted directly from the original dataset.

[0082] The construction module 102 is used to construct an individual interpersonal contact number inference model: based on the individual demographic characteristics and individual travel activity characteristics extracted from a small sample questionnaire survey, a machine learning inference model for the number of individual interpersonal contacts is constructed.

[0083] In this embodiment, each community resident extracted from a small-sample questionnaire survey is considered a sample, and machine learning methods are used to solve the problem of inferring individual interpersonal contact patterns. First, the samples are divided into training and testing sets, and individual demographic characteristics and individual travel activity characteristics are used as input variables to train the machine learning model. Based on this, an inference model for individual interpersonal contact patterns is constructed to infer the number of contacts an individual has with people of different age groups in different scenarios. Through this inference model, the individual's contact patterns are accurately inferred, providing important data support for subsequent research on regional interpersonal contact patterns. Specifically: The construction module 102 defines the model: This embodiment uses a machine learning model to infer the mean and median of individual interpersonal contact patterns, and selects the specific model with better performance. The specific definition of the model is as follows: Random Forest Model: Random forest is an ensemble learning method that enhances the inference accuracy of a model by constructing multiple decision trees and combining their inference results. Each tree is trained based on a random subset of samples and a random subset of features. Random forests have strong resistance to overfitting and achieve good results in both classification and regression tasks.

[0084] XGBoost (eXtreme Gradient Boosting) model: XGBoost is a scalable gradient boosting decision tree algorithm. It effectively prevents overfitting by introducing a regularization term into the objective function to control model complexity. XGBoost performs a second-order Taylor expansion of the objective function and supports feature parallelism and data parallelism, resulting in excellent performance in both accuracy and computational efficiency.

[0085] LightGBM (Light Gradient Boosting Machine) model: LightGBM is an efficient gradient boosting framework. It employs a histogram-based decision tree algorithm and introduces two techniques: gradient one-sided sampling (GOSS) and mutually exclusive feature binding (EFB). This significantly improves training speed and reduces memory consumption while maintaining accuracy, making it particularly suitable for large-scale datasets.

[0086] CatBoost (Categorical Boosting) model: CatBoost is a machine learning algorithm based on gradient boosting decision trees. Its core advantage lies in its ability to efficiently and directly process categorical features without extensive preprocessing. By using key techniques such as ordered boosting and target statistics, it effectively avoids gradient bias and prediction offset, demonstrating excellent performance on datasets containing a large number of categorical features.

[0087] In other embodiments, the machine learning algorithm can be changed from the random forest algorithm to other machine learning or deep learning algorithms, such as support vector regression, feedforward neural networks, and elastic networks.

[0088] Table 3 Model List

[0089] The construction module 102 performs parameter settings and optimization: This embodiment performs model parameter tuning in two stages: First, the default parameters of each machine learning model are used as a baseline; then, a random cross-reference method is used to optimize hyperparameters to improve model performance. Different loss functions are tested on the mean prediction model to better infer the right-skewed data distribution.

[0090] Table 4. List of Loss Functions for Mean Prediction Model

[0091] Randomized SearchCV (Randomized Crossover Search): Randomized crossover search is an efficient hyperparameter optimization algorithm. Unlike grid search, which traverses all possible combinations, Randomized crossover search randomly selects and evaluates a fixed number of parameter combinations from a specified parameter distribution. It can discover high-performance parameter combinations with a higher probability and in a shorter time, making it particularly suitable for high-dimensional parameter spaces. Its basic process is as follows: Parameter space: Specifies a probability distribution (such as uniform distribution, log-uniform distribution, etc.) for each hyperparameter, rather than a discrete list of values.

[0092] Random sampling: randomly selecting from this parameter space Group hyperparameter configuration, among which The pre-set number of iterations (calculation budget).

[0093] Evaluation and selection: Evaluate the model performance corresponding to each set of parameter configurations on the validation set, and finally select the hyperparameter combination with the best performance.

[0094] This process can be formally represented as finding the optimal hyperparameter configuration. :

[0095] in: This represents a set of hyperparameter configurations.

[0096] It is a predefined joint probability distribution of hyperparameters.

[0097] Is it using hyperparameter configuration? The performance evaluation function of the model on the validation set (such as the coefficient of determination) ).

[0098] It represents the total number of random samples (for budget calculation).

[0099] This indicates the search for a performance evaluation function. Maximize hyperparameter configuration .

[0100] The evaluation metrics defined in the construction module 102 are as follows: This embodiment uses the coefficient of determination (R²), mean absolute percentage error (MAPE), mean absolute error (MAE), and root mean square error (RMSE) as indicators to evaluate the inference performance and error of the model, and their specific definitions are as follows: Table 5. List of Model Evaluation Indicators

[0101] The inference module 103 is used for inferring regional interpersonal contact patterns based on big data. That is, under the premise of privacy protection, representative individual samples and their individual characteristics are extracted from mobile phone location big data for each region, input into a trained model, and large-scale individual interpersonal contact patterns with regional representativeness are inferred. Finally, the individual inference results are aggregated to the regional level to obtain the interpersonal contact patterns of the entire region. Specifically: The inference module 103 extracts representative samples: This embodiment uses Shenzhen as an example to illustrate the process. Representative individual samples of different age groups in various regions are extracted from mobile phone big data to ensure that these samples cover all age groups and match the actual age distribution of the regional population. For infants and toddlers aged 0-6, whose mobile phone usage is relatively low, individual samples from small-sample questionnaires are used to supplement the data and compensate for any missing information. Ultimately, 1.66 million mobile phone users were extracted.

[0102] The inference module 103 performs regional interpersonal contact pattern inference: This embodiment inputs age-representative individual contact data within a region into a trained machine learning model to infer interpersonal contact patterns across the entire region. Based on the output of the machine learning model, mean and median statistics reflecting the contact characteristics of the population in the region are generated, providing a crucial data foundation for the analysis and modeling of infectious disease transmission dynamics.

[0103] The inference module 103 performs regional interpersonal contact pattern inference and evaluation: Please also refer to Figure 3 This embodiment constructs an SEIR transmission model for typical respiratory infectious diseases (such as influenza and COVID-19). By comparing the simulation results under different contact number parameters, it analyzes the impact of contact heterogeneity on transmission dynamics and compares it with the traditional method that assumes uniform contact patterns, thus verifying the advantages of this application and improving the generalizability and reliability of the transmission model in different regions.

[0104] This invention utilizes small-sample questionnaire survey data and mobile phone location big data to extract consistent individual demographic and travel activity characteristics from these "large" and "small" data sets. It then constructs a model for inferring the number of interpersonal contacts by integrating demographic and travel activity characteristics, thereby enabling the inference of fine-grained interpersonal contact patterns within cities. This invention not only improves the accuracy of interpersonal contact pattern inference, providing strong data support for developing more precise public health interventions, but also provides a solid foundation and practical example for future inference of interpersonal contact patterns in different regions.

[0105] Although the present invention has been described with reference to the present preferred embodiments, those skilled in the art should understand that the above preferred embodiments are only used to illustrate the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An urban intra-zonal interpersonal contact pattern inference method characterized by comprising: The method includes: Step S1: Design a small-sample questionnaire survey covering demographic and travel activity characteristics, and extract individual demographic and travel activity characteristics from it; Step S2: Based on the individual demographic characteristics and individual travel activity characteristics extracted from the small sample questionnaire survey, construct a machine learning inference model for the number of interpersonal contacts of an individual. Step S3: Under the premise of privacy protection, extract representative individual samples and their individual characteristics from mobile phone location big data, input them into the trained inference model, infer large-scale individual interpersonal contact patterns with regional representativeness, and finally aggregate the individual inference results to the regional level to obtain the interpersonal contact patterns of the entire region.

2. The method of claim 1, wherein, Step S1 includes: Step S11: Extract individual demographic characteristics through small-sample questionnaire surveys and mobile phone location big data. Step S12: Based on the individual's travel location on the same day, obtain the characteristics of the individual's travel activities through the survey data.

3. The method as described in claim 2, characterized in that, Step S2 includes: Step S21: Use machine learning models to infer the mean and median of individual interpersonal contact patterns, and select the specific model with better performance. Step S22: Use the default parameters of each machine learning model as a baseline; then use a random cross-search method to optimize the hyperparameters. Step S23: The coefficient of determination, mean absolute percentage error, mean absolute error, and root mean square error are used as indicators to evaluate the inference performance and error of the model.

4. The method as described in claim 3, characterized in that, Step S3 includes: Step S31: Extract representative individual samples of different age groups in each region from mobile phone big data to ensure that the representative individual samples can cover people of all ages and match the actual age distribution of the population in the region. Step S32: Input age-representative individual contact data within the region into the trained machine learning model to infer the interpersonal contact patterns of the entire region; based on the output of the machine learning model, generate mean and median statistics reflecting the contact characteristics of the population in the region.

5. The method as described in claim 4, characterized in that, Step S3 further includes: Step S33: Compare the simulation results under different contact number parameters, analyze the impact of contact heterogeneity on propagation dynamics, and compare them with the traditional method that assumes uniform contact patterns, so as to improve the generalizability and reliability of the propagation model in different regions.

6. A system for inferring interpersonal contact patterns in urban areas, characterized in that, The system includes an extraction module, a construction module, and an inference module, wherein: The extraction module is used to design a small-sample questionnaire survey covering demographic and travel activity characteristics, and to extract individual demographic and travel activity characteristics from it. The construction module is used to build a machine learning inference model for the number of interpersonal contacts of an individual based on individual demographic characteristics and individual travel activity characteristics extracted from a small sample questionnaire survey. The inference module is used to extract representative individual samples and their individual characteristics from mobile phone location big data under the premise of privacy protection, input them into the trained inference model, infer large-scale individual interpersonal contact patterns with regional representativeness, and finally aggregate the individual inference results to the regional level to obtain the interpersonal contact patterns of the entire region.

7. The system as described in claim 6, characterized in that, The extraction module is specifically used for: Individual demographic characteristics were extracted through small-sample questionnaire surveys and mobile phone location big data. By using survey data, we can obtain the characteristics of an individual's travel activities based on their location on the same day.

8. The system as described in claim 7, characterized in that, The aforementioned building module is specifically used for: Machine learning models are used to infer the mean and median of individual interpersonal contact patterns, and specific models with better performance are selected from them. The default parameters of each machine learning model were used as a baseline; then, a random cross-search method was used for hyperparameter optimization. The coefficient of determination, mean absolute percentage error, mean absolute error, and root mean square error are used as indicators to evaluate the inference performance and error of the model.

9. The system as described in claim 8, characterized in that, The inference module is specifically used for: Representative individual samples of different age groups in various regions are extracted from mobile phone big data to ensure that the representative individual samples can cover people of all ages and match the actual age distribution of the population in the region. The contact data of age-representative individuals within the region are input into a trained machine learning model to infer interpersonal contact patterns across the entire region. Based on the output of the machine learning model, mean and median statistics reflecting the contact characteristics of the population in the region are generated.

10. The system as described in claim 9, characterized in that, The inference module is also specifically used for: The simulation results under different contact number parameters were compared to analyze the impact of contact heterogeneity on the propagation dynamics. The results were also compared with those of traditional methods that assume uniform contact patterns, in order to improve the generalizability and reliability of the propagation model in different regions.