Method and system for generating a screening model, screening a high-risk population for infectious disease
By generating a screening model for high-risk groups of infectious diseases, and utilizing user trajectory information and machine learning algorithms, the problem of rapidly screening high-risk groups of infectious diseases has been solved, improving the efficiency of infectious disease control and saving human resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- THE FOURTH PARADIGM BEIJING TECH CO LTD
- Filing Date
- 2020-04-22
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies are insufficient for quickly and accurately screening high-risk groups for infectious diseases, resulting in low efficiency in infectious disease control.
By acquiring user trajectory information, establishing a sample table and extracting features, a screening model for high-risk groups of infectious diseases is generated using machine learning algorithms. The machine learning model is trained using users' mobile terminal data to predict the degree of risk of users contracting infectious diseases.
It enables rapid and accurate screening of high-risk groups for infectious diseases, improving the effectiveness of infectious disease control and saving human resources.
Smart Images

Figure CN115171910B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application filed on April 22, 2020, with application number 202010323348.X, entitled "Generation of screening model, method and system for screening high-risk populations for infectious diseases". Technical Field
[0002] This invention generally relates to the field of artificial intelligence, and more specifically, to a method and system for generating a screening model for high-risk groups of infectious diseases, and a method and system for screening high-risk groups of infectious diseases. Background Technology
[0003] The primary transmission route for some infectious diseases is close-range human-to-human transmission. Therefore, quickly and accurately identifying and monitoring high-risk groups is one of the most effective means of controlling the spread of infectious diseases. Summary of the Invention
[0004] An exemplary embodiment of the present invention provides a method and system for generating screening models and screening high-risk groups for infectious diseases, which can be used to quickly and accurately screen high-risk groups for a certain infectious disease.
[0005] According to an exemplary embodiment of the present invention, a method for generating a screening model for high-risk groups of infectious diseases is provided, wherein the method includes: acquiring a training dataset, wherein the training dataset includes user trajectory information, wherein the user trajectory information is obtained based on user mobile terminal related data; establishing a sample table, wherein each sample in the sample table includes a user identifier and a sample label, the sample label indicating whether the user is a positive sample user diagnosed with / suspected infection with a specified type of infectious disease or a normal negative sample user; extracting features for each sample in the sample table based on the training dataset, and incorporating the extracted features into the sample table; and using a machine learning algorithm to train a machine learning model based on the sample table with the incorporated features, thereby generating a screening model for high-risk groups of the specified type of infectious disease.
[0006] Optionally, the user trajectory information includes: the base station identifier of the base station used by the user's mobile terminal in each time period, wherein each time period is obtained by dividing a specific time span according to a preset time granularity.
[0007] Optionally, the base station used by the user's mobile terminal in each time period is: the base station used by the user's mobile terminal for the longest time in each time period or the base station used by the user's mobile terminal at a specified time point in each time period.
[0008] Optionally, the training dataset includes at least one of the following tables: a table of confirmed users, including the user IDs of positive sample users and their diagnosis times; a user trajectory table, including the user IDs and the base station IDs of the base stations used by the user's mobile terminal in each time period; a user information table, including the user IDs and the user's attribute information; a base station table, including the base station IDs and the geographical location information of the base stations; a user communication record table, including the user IDs and the communication records between the user and other users of mobile terminals using the user's mobile terminal; and a user address book information table, including the user IDs and the user IDs of contacts in the address book of at least one application of the user's mobile terminal.
[0009] Optionally, the step of extracting features for each sample in the sample table based on the training dataset includes: directly processing the information in the data table included in the training dataset into basic features corresponding to each user ID; and / or generating derived features corresponding to each user ID based on the information in the data table included in the training dataset, wherein the derived features include at least one of the following: aggregated features regarding the user's activity level, aggregated features regarding the social intimacy between the user and positive sample users, aggregated features regarding the user and positive sample users appearing in the same base station area during the same time period, and aggregated features regarding the user appearing in a susceptible area, wherein the base station area where the user appears during a time period is the area corresponding to the base station used by the user's mobile terminal during that time period, and the susceptible area is the base station area where each positive sample user has appeared in each time period.
[0010] Optionally, the aggregated features regarding a user's activity level include at least one of the following: features indicating the number of all base station areas the user had visited before a specific time, features indicating the number of all provinces / cities the user had visited before a specific time, features indicating the maximum longitude of all base station areas the user had visited before a specific time, features indicating the minimum longitude of all base station areas the user had visited before a specific time, features indicating the maximum latitude of all base station areas the user had visited before a specific time, features indicating the minimum latitude of all base station areas the user had visited before a specific time, features indicating the user's activity distance before a specific time, and features indicating whether the user is from another province; the aggregated features regarding the user's social intimacy with positive sample users include: features indicating different social intimacy levels with the user. The statistical characteristics of the number of positive sample users, wherein the social intimacy between users is determined based on call records and / or user contact information; the aggregated characteristics of users appearing in the same base station area with positive sample users in the same time period include: the statistical characteristics of the number of times the user and positive sample users appeared in the same base station area in the same time period; the aggregated characteristics of users appearing in susceptible areas include at least one of the following: the statistical characteristics of the number of times the user appeared in each susceptible area, and the statistical characteristics of the number of times the user appeared in susceptible areas of different risk levels, wherein the risk level of susceptible areas is related to the number of times positive sample users have appeared, wherein if the user is a positive sample user, the specific time corresponding to the user is the time of the user's diagnosis; if the user is a negative sample user, the specific time corresponding to the user is the time assigned to the user according to a specific rule.
[0011] Optionally, the step of generating a feature of the statistical value regarding the number of times a user and a positive sample user appear in the same base station area during the same time period includes: constructing a dictionary of all activity trajectories of all positive sample users before their diagnosis time based on the user trajectory table and the confirmed user table, wherein each element in the dictionary is a trajectory point representing a positive sample user appearing in a base station area during a time period; for each user, determining whether the user's trajectory point overlaps with the trajectory point of a positive sample user in the dictionary before a specific time, and statistically analyzing the overlapping trajectory points to obtain a feature of the statistical value regarding the number of times the user and a positive sample user appear in the same base station area during the same time period.
[0012] Optionally, the user information table includes at least one of the following user attribute information: package tariff, package data usage, package call duration, package SMS message count, monthly internet data usage, monthly call duration, monthly SMS message count, monthly call charge, average monthly call duration, average monthly internet data usage, average monthly SMS message count, average monthly call charge, whether it is a corporate user, mobile phone number registration location, number of years of network access, age, and gender.
[0013] Optionally, the step of obtaining the training dataset includes: obtaining signaling data of communication between the user's mobile terminal and the base station within the specific time span; and obtaining the base station ID of the base station used by each user's mobile terminal in each time period of the specific time span based on the signaling data.
[0014] Optionally, the machine learning algorithm is a model fusion algorithm.
[0015] Optionally, the specific time assigned to the user according to the specific rule is the last day of the specific time span, or the specific time corresponding to each of the negative sample users is uniformly set according to the distribution of the diagnosis time of all positive sample users.
[0016] According to another exemplary embodiment of the present invention, a method for screening high-risk groups for infectious diseases is provided, wherein the method includes: acquiring a prediction dataset of users to be screened, wherein the prediction dataset includes trajectory information of the users to be screened, wherein the trajectory information of the users to be screened is obtained based on mobile terminal-related data of the users to be screened; extracting features for each user to be screened based on the prediction dataset; using a high-risk group screening model for a specified type of infectious disease generated by performing the method described above for generating a high-risk group screening model for infectious diseases, predicting the risk level of the user to be screened from contracting the specified type of infectious disease based on the extracted features; and outputting the predicted risk level of the user from contracting the specified type of infectious disease.
[0017] Optionally, the step of outputting the predicted risk level of a user being infected with the specified type of infectious disease includes: outputting the ranking results of users in descending order of predicted risk level; and / or, outputting only the risk level of users whose predicted risk level meets preset conditions.
[0018] According to another exemplary embodiment of the present invention, a system for generating a screening model for high-risk groups of infectious diseases is provided, wherein the system comprises: a dataset acquisition device adapted to acquire a training dataset, wherein the training dataset includes user trajectory information, wherein the user trajectory information is obtained based on user mobile terminal related data; a sample table building device adapted to build a sample table, wherein each sample in the sample table includes a user identifier and a sample label, the sample label indicating whether the user is a positive sample user diagnosed with / suspected infection with a specified type of infectious disease or a normal negative sample user; a feature extraction device adapted to extract features for each sample in the sample table based on the training dataset, and incorporate the extracted features into the sample table; and a training device adapted to use a machine learning algorithm to train a machine learning model based on the sample table with incorporated features, thereby generating a screening model for high-risk groups of the specified type of infectious disease.
[0019] Optionally, the user trajectory information includes: the base station identifier of the base station used by the user's mobile terminal in each time period, wherein each time period is obtained by dividing a specific time span according to a preset time granularity.
[0020] Optionally, the base station used by the user's mobile terminal in each time period is: the base station used by the user's mobile terminal for the longest time in each time period or the base station used by the user's mobile terminal at a specified time point in each time period.
[0021] Optionally, the training dataset includes at least one of the following tables: a table of confirmed users, including the user IDs of positive sample users and their diagnosis times; a user trajectory table, including the user IDs and the base station IDs of the base stations used by the user's mobile terminal in each time period; a user information table, including the user IDs and the user's attribute information; a base station table, including the base station IDs and the geographical location information of the base stations; a user communication record table, including the user IDs and the communication records between the user and other users of mobile terminals using the user's mobile terminal; and a user address book information table, including the user IDs and the user IDs of contacts in the address book of at least one application of the user's mobile terminal.
[0022] Optionally, the feature extraction device is adapted to directly process the information in the data table included in the training dataset into basic features corresponding to each user ID; and / or, generate derived features corresponding to each user ID based on the information in the data table included in the training dataset, wherein the derived features include at least one of the following: aggregated features regarding the user's activity level, aggregated features regarding the social intimacy between the user and positive sample users, aggregated features regarding the user and positive sample users appearing in the same base station area during the same time period, and aggregated features regarding the user appearing in a susceptible area, wherein the base station area where the user appears during a time period is the area corresponding to the base station used by the user's mobile terminal during that time period, and the susceptible area is the base station area where each positive sample user has appeared in each time period.
[0023] Optionally, the aggregated features regarding a user's activity level include at least one of the following: features indicating the number of all base station areas the user had visited before a specific time, features indicating the number of all provinces / cities the user had visited before a specific time, features indicating the maximum longitude of all base station areas the user had visited before a specific time, features indicating the minimum longitude of all base station areas the user had visited before a specific time, features indicating the maximum latitude of all base station areas the user had visited before a specific time, features indicating the minimum latitude of all base station areas the user had visited before a specific time, features indicating the user's activity distance before a specific time, and features indicating whether the user is from another province; the aggregated features regarding the user's social intimacy with positive sample users include: features indicating different social intimacy levels with the user. The statistical characteristics of the number of positive sample users, wherein the social intimacy between users is determined based on call records and / or user contact information; the aggregated characteristics of users appearing in the same base station area with positive sample users in the same time period include: the statistical characteristics of the number of times the user and positive sample users appeared in the same base station area in the same time period; the aggregated characteristics of users appearing in susceptible areas include at least one of the following: the statistical characteristics of the number of times the user appeared in each susceptible area, and the statistical characteristics of the number of times the user appeared in susceptible areas of different risk levels, wherein the risk level of susceptible areas is related to the number of times positive sample users have appeared, wherein if the user is a positive sample user, the specific time corresponding to the user is the time of the user's diagnosis; if the user is a negative sample user, the specific time corresponding to the user is the time assigned to the user according to a specific rule.
[0024] Optionally, the feature extraction device is adapted to construct a dictionary of all activity trajectories of all positive sample users before their diagnosis time based on the user trajectory table and the confirmed user table; for each user, it determines whether the user's trajectory points overlap with the trajectory points of positive sample users in the dictionary whose diagnosis time was before a specific time, and performs statistics on the overlapping trajectory points to obtain a feature of the statistical value of the number of times the user and positive sample users appeared in the same base station area in the same time period, wherein each element in the dictionary is a trajectory point used to characterize a positive sample user appearing in a base station area in a time period.
[0025] Optionally, the user information table includes at least one of the following user attribute information: package tariff, package data usage, package call duration, package SMS message count, monthly internet data usage, monthly call duration, monthly SMS message count, monthly call charge, average monthly call duration, average monthly internet data usage, average monthly SMS message count, average monthly call charge, whether it is a corporate user, mobile phone number registration location, number of years of network access, age, and gender.
[0026] Optionally, the dataset acquisition device is adapted to acquire signaling data of communication between a user's mobile terminal and a base station within the specific time span; and based on the signaling data, acquire the base station ID of the base station used by each user's mobile terminal in each time period of the specific time span.
[0027] Optionally, the machine learning algorithm is a model fusion algorithm.
[0028] Optionally, the specific time assigned to the user according to the specific rule is the last day of the specific time span, or the specific time corresponding to each of the negative sample users is uniformly set according to the distribution of the diagnosis time of all positive sample users.
[0029] According to another exemplary embodiment of the present invention, a system for screening high-risk groups for infectious diseases is provided, wherein the system comprises: a dataset acquisition device adapted to acquire a predicted dataset of users to be screened, wherein the predicted dataset includes trajectory information of the users to be screened, wherein the trajectory information of the users to be screened is obtained based on mobile terminal-related data of the users to be screened; a feature extraction device adapted to extract features for each user to be screened based on the predicted dataset; a prediction device adapted to predict the risk level of the user to be screened for the specified type of infectious disease based on the extracted features using a high-risk group screening model for a specified type of infectious disease generated by the system for generating high-risk groups for infectious diseases as described above; and an output device adapted to output the predicted risk level of the user for the specified type of infectious disease.
[0030] Optionally, the output device is adapted to output the user's ranking results in descending order of predicted risk level; and / or, the output device is adapted to output only the risk level of users whose predicted risk level meets preset conditions.
[0031] According to another exemplary embodiment of the present invention, a system is provided that includes a storage device comprising at least one computing device and at least one storage instruction, wherein, when the instruction is executed by the at least one computing device, it causes the at least one computing device to perform the method described above for generating a screening model for high-risk groups of infectious diseases and / or the method described above for screening high-risk groups of infectious diseases.
[0032] According to another exemplary embodiment of the present invention, a computer-readable storage medium storing instructions is provided, wherein when the instructions are executed by at least one computing device, the at least one computing device causes the at least one computing device to perform the method described above for generating a screening model for high-risk groups of infectious diseases and / or the method described above for screening high-risk groups of infectious diseases.
[0033] The method and system for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention can generate a predictive model for predicting the risk of a user contracting a certain infectious disease based on the activity trajectory of a mobile terminal user. Furthermore, in addition to generating basic features based on multi-dimensional data, a large number of derived features are constructed, thereby improving the predictive performance of the generated screening model. The method and system for screening high-risk groups of infectious diseases according to an exemplary embodiment of the present invention can quickly, effectively, and labor-savingly predict the risk of a mobile terminal user contracting a certain infectious disease.
[0034] Further aspects and / or advantages of the general concept of the invention will be set forth in part in the description which follows, and in part will be obvious from the description or may be learned by practice of the general concept of the invention. Attached Figure Description
[0035] The above and other objects and features of exemplary embodiments of the present invention will become clearer from the following description taken in conjunction with the accompanying drawings, which illustrate exemplary embodiments, wherein:
[0036] Figure 1 A flowchart illustrating a method for generating a screening model for high-risk populations of infectious diseases according to an exemplary embodiment of the present invention;
[0037] Figure 2 A flowchart illustrating a method for screening high-risk populations for infectious diseases according to an exemplary embodiment of the present invention;
[0038] Figure 3A block diagram of a system for generating a screening model for high-risk populations of infectious diseases according to an exemplary embodiment of the present invention is shown.
[0039] Figure 4 A block diagram of a system for screening high-risk populations for infectious diseases according to an exemplary embodiment of the present invention is shown. Detailed Implementation
[0040] The present invention will now be described in detail with reference to embodiments thereof, examples of which are illustrated in the accompanying drawings, wherein the same reference numerals refer to the same parts throughout. The embodiments will be described below with reference to the accompanying drawings in order to explain the present invention.
[0041] Figure 1 A flowchart illustrating a method for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention is provided. The screening model for high-risk groups of infectious diseases generated by the method can predict the risk level of a mobile terminal user contracting a specified type of infectious disease. For example, the specified type of infectious disease may be an infectious disease whose primary transmission route is close-range person-to-person contact.
[0042] Reference Figure 1 In step S10, a training dataset is obtained, wherein the training dataset includes user trajectory information, which is obtained based on user mobile terminal related data.
[0043] As an example, the user trajectory information is information that can reflect the user's activity trajectory, such as information that can reflect the user's location at various points in time.
[0044] As an example, the user's mobile terminal-related data may be data that can be used to determine the location of the mobile terminal. For example, the user's mobile terminal-related data may include: the mobile terminal's location data and / or communication data, such as signaling data between the user's mobile terminal and a base station.
[0045] In addition, as an example, the training dataset may also include other information obtained based on the user's mobile terminal-related data. For example, the training dataset may also include the user's communication records with other users of mobile terminals and / or the user's mobile terminal's address book information.
[0046] As an example, the user trajectory information may include: the base station identifier of the base station used by the user's mobile terminal in each time period, wherein each time period is obtained by dividing a specific time span according to a preset time granularity.
[0047] As an example, the base station used by a user's mobile terminal in each time period can refer to: the base station used by the user's mobile terminal for the longest time in each time period, or the base station used by the user's mobile terminal at a specific time point in each time period. For example, the specific time point in each time period can be an end point and / or a certain time point in the middle of that time period.
[0048] The specific time span refers to the time span corresponding to the training data. For example, the specific time span can be a certain date range. As an example, the preset time granularity can be set according to actual conditions and needs. For example, the preset time granularity can be 15 minutes. For example, the training dataset can include the base station identifier of the base station used by the user's mobile terminal for the longest time within each 15 minutes. In fact, this information can reflect the area where the user of the mobile terminal spent the most time within each 15 minutes (that is, the base station area mentioned later), thereby reflecting the user's activity trajectory within the specific time span to a certain extent. For example, the training dataset can include: the base station identifier of the base station used by the user's mobile terminal every preset time granularity (e.g., every 15 minutes).
[0049] As an example, signaling data of communication between a user's mobile terminal and a base station within the specific time span can be obtained; and based on the signaling data, the base station ID of the base station used by each user's mobile terminal in each time period within the specific time span can be obtained. For example, based on the signaling data, the base stations used by the user's mobile terminal in each time period and the duration of use can be determined, thereby determining the base station ID of the base station used for the longest time in each time period. For example, based on the signaling data of communication between the user's mobile terminal and the base station at each preset time granularity, the base station ID of the base station used by the user's mobile terminal at a specified time point within each time period can be determined.
[0050] As an example, the training dataset may include at least one of the following data tables: a table of diagnosed users, a user trajectory table, a user information table, a base station table, a user communication record table, and a user address book information table. It should be understood that the training dataset may also include other types of data tables, and this invention is not limited thereto. Examples of each data table will be provided in detail later.
[0051] In step S20, a sample table is created, wherein each sample in the sample table includes a user identifier and a sample tag.
[0052] Here, the sample label indicates whether the user is a positive sample user who has been diagnosed with or is suspected of being infected with the specified type of infectious disease, or a normal negative sample user. A normal negative sample user is a user who has not been diagnosed with or is suspected of being infected with the specified type of infectious disease.
[0053] As an example, the user identifier could be a user id.
[0054] As an example, a sample table can be built based on the data tables included in the training dataset. For instance, a sample table can be built based on a table of diagnosed users and a table of user trajectories.
[0055] In step S30, based on the training dataset, features are extracted for each sample in the sample table, and the extracted features are incorporated into the sample table.
[0056] As an example, the information in the data tables included in the training dataset can be directly processed into basic features corresponding to each user ID; and / or, derived features corresponding to each user ID can be generated based on the information in the data tables included in the training dataset. That is, the extracted features may include basic features and / or derived features. Examples of extracted basic and derived features will be provided in detail later.
[0057] Furthermore, as an example, the extracted basic features and / or derived features can be concatenated to the corresponding samples based on the user ID to obtain a sample table that incorporates the features.
[0058] In step S40, a machine learning algorithm is used to train a machine learning model based on a sample table incorporating features, thereby generating a screening model for high-risk groups of the specified type of infectious disease.
[0059] It should be understood that various suitable machine learning algorithms can be used to train the machine learning model. As an example, the machine learning algorithm could be a model fusion (stacking) algorithm.
[0060] Model fusion algorithms are a hierarchical model integration framework. Taking a two-layer model as an example, the first layer consists of multiple base learners whose input is a sample table. The second layer model is retrained using the output of the first layer's base learners as the training set to obtain a complete stacking model. In other words, multiple different models are selected for prediction, and then a prediction model is built on top of the prediction results of each model to obtain the final prediction result. Model fusion can combine the advantages of different models to further improve model performance. For example, the model fusion algorithm used in the exemplary embodiments of this invention can employ algorithms such as logistic regression and GBDT (Gradient Boosting Decision Tree) for model fusion.
[0061] The following will provide detailed examples illustrating the data tables included in the training dataset.
[0062] Specifically, the confirmed user table may include the user ID of the positive sample users and their diagnosis time.
[0063] The user trajectory table is a time-series table, which can include the user ID and the base station ID of the base station used by the user's mobile terminal in each time period, and the user ID is the primary key of the table.
[0064] The user information table is a static table and may include user ID and user attribute information. As an example, the attribute information may include at least one of the following: package tariff, package data allowance, package call duration, package SMS message count, monthly internet data usage (e.g., internet data usage for each month within the specific time span), monthly call duration, monthly SMS message count, monthly call charge, average monthly call duration (MOU) (e.g., the average of all months within the specific time span), average monthly internet data usage (DOU), average monthly SMS message count, average monthly call charge (ARPU), whether it is a group user, mobile phone number registration location (e.g., down to province, city, district, and county), number of years of service, age, and gender. The user ID is the primary key of the table.
[0065] The base station table is a static table and may include the base station ID and the geographical location information of the base station. For example, the geographical location information of the base station may include at least one of the following: the longitude of the base station, the latitude of the base station, and the province, city, district, and county where the base station is located. Among them, the base station ID is the primary key of the table.
[0066] The user communication record table may include the user ID and communication records between the user and other users using the mobile terminal. For example, communication here can refer to communication methods that can be recorded by the base station, such as telephone communication, SMS communication, and internet communication. As an example, communication records between the user and other users using the mobile terminal can be obtained based on the signaling data of the communication between the user's mobile terminal and the base station.
[0067] The user contact list information table may include the user ID and the user IDs of contacts in the contact list of at least one application on the user's mobile device. For example, the at least one application may include a dialer application and / or an instant messaging application.
[0068] As an example, the basic features derived by directly processing the information from the data tables included in the training dataset may include at least one of the following: features indicating age, features indicating gender, features indicating monthly call charges, features indicating average monthly call duration (MOU), features indicating average monthly internet data usage (DOU), features indicating network duration based on years of service, and features indicating average monthly call charges per user (ARPU). It should be understood that the basic features may also include other features that can be obtained by directly processing the information from the data tables included in the training dataset.
[0069] The following will provide detailed examples illustrating the derived features generated for each sample.
[0070] As an example, the derived features may include at least one of the following: aggregated features regarding the user's activity level, aggregated features regarding the user's social intimacy with positive sample users, aggregated features regarding the user and positive sample users appearing in the same base station area during the same time period, and aggregated features regarding the user appearing in a susceptible area. It should be understood that the derived features may also include other types of features that can be obtained based on the user's mobile terminal-related data.
[0071] Here, the base station area where a user appears within a certain time period refers to the area corresponding to the base station used by the user's mobile terminal during that time period. In other words, the user stayed in that base station area during that time period, and thus, to a certain extent, the base station area can be used to characterize the user's activity location during that time period. For example, the area corresponding to a base station can be understood as the area where the user's mobile terminal will use the base station when it enters that area.
[0072] The susceptible region is the base station area where each positive sample user appeared at each time period. In other words, it is the base station area where a positive sample user appeared at a certain time period. In other words, the set of base station areas where all positive sample users appeared at each time period is the susceptible region.
[0073] As an example, aggregated features regarding a user's activity level may include at least one of the following: a feature indicating the number of all base station areas the user has visited in the n days prior to a specific time; a feature indicating the number of all provinces / cities the user has visited in the n days prior to a specific time; a feature indicating the maximum longitude of all base station areas the user has visited in the n days prior to a specific time; a feature indicating the minimum longitude of all base station areas the user has visited in the n days prior to a specific time; a feature indicating the maximum latitude of all base station areas the user has visited in the n days prior to a specific time; a feature indicating the minimum latitude of all base station areas the user has visited in the n days prior to a specific time; a feature indicating the maximum activity distance (e.g., lateral distance / vertical distance / Euclidean distance, etc.) of the user in the n days prior to a specific time; and a feature indicating whether the user is from another province. Here, n is an integer greater than 0, and for example, n can be a series of values, such as 1, 2, ..., 7.
[0074] As an example, if the user is a positive sample user, the specific time corresponding to them can be the time of their diagnosis; if the user is a negative sample user, the specific time corresponding to them can be the time assigned to them according to a specific rule. For example, the specific time assigned to the user according to the specific rule can be the last day of the specific time span, or, the specific time corresponding to all negative sample users can be uniformly set according to the distribution of the diagnosis times of all positive sample users. For example, the distribution of the specific times corresponding to all negative sample users can be made consistent with the distribution of the diagnosis times of all positive sample users.
[0075] As an example, whether a user is from another province can be determined by whether their mobile phone number registration location matches their primary activity location. As an example, the longitude and latitude of a base station area can be determined based on the base station's geographical location information, thereby identifying the maximum and minimum longitude, maximum and minimum latitude, maximum activity distance, and primary activity location of all base station areas where the user has appeared. As an example, the province / city where the base station area is located can be determined based on the base station's geographical location information, thereby identifying all provinces / cities where the user has appeared.
[0076] As an example, n can be 1, 2, ..., 7 in sequence. For example, the features used to indicate the number of all base station areas that the user has appeared in within n days before a specific time may include: features used to indicate the total number of different base station areas that the user has appeared in during various time periods within 1 day before a specific time, features used to indicate the total number of different base station areas that the user has appeared in during various time periods within 2 days before a specific time, ..., features used to indicate the total number of different base station areas that the user has appeared in during various time periods within 7 days before a specific time.
[0077] As an example, aggregated features regarding the social intimacy between a user and positive sample users may include: features regarding the statistical value of the number of positive sample users with different social intimacy with that user, wherein the social intimacy between users is determined based on call records between users and / or the user's address book information.
[0078] As an example, social intimacy between users can be characterized by the shortest social distance between them, and the shorter the shortest social distance, the higher the social intimacy between users. As an example, the characteristics of the statistical value regarding the number of positive sample users with different social intimacy with a given user may include at least one of the following: a characteristic indicating the number of positive sample users with a shortest social distance of i with that user, and a characteristic indicating the proportion of the number of positive sample users with a shortest social distance of i with that user to the total number of users with a shortest social distance of i with that user, where i is an integer greater than 0, and i can be multiple values, for example, 1, 2, ..., 7.
[0079] As an example, a social distance graph can be constructed based on communication records between users and / or other users recorded in a user's address book, to determine the shortest social distance between a user and other users. For example, if users have direct communication records or are recorded in each other's address books, i.e., they are directly adjacent, the shortest social distance between them can be defined as 1. If users are not directly adjacent but have a social relationship indirectly through another user, the shortest social distance between them can be defined as 2.
[0080] In addition to determining the social intimacy between users based on their direct communication records with other users, and the communication records between users who have communication records with the user (directly or indirectly through users who have communication records with the user) and other users, the social intimacy between users can also be determined based on the number of communications between users, with the higher the number of communications, the higher the social intimacy.
[0081] According to an embodiment of the present invention, considering that infectious diseases are very likely to spread among users with high social intimacy (e.g., relatives, colleagues, friends), aggregated features of social intimacy between users and positive sample users are used in the model to improve the predictive performance of the model.
[0082] As an example, the aggregated features of a user's presence in a vulnerable area may include at least one of the following: a feature of the statistical value of the number of times the user appears in each vulnerable area, and a feature of the statistical value of the number of times the user appears in vulnerable areas of different risk levels, wherein the risk level of a vulnerable area is related to the number of times positive sample users have appeared, and the more times all positive sample users have appeared, the higher the risk level of the vulnerable area.
[0083] As an example, the characteristics of the statistical value regarding the number of times the user appears in each susceptible region may include at least one of the following: a characteristic indicating the total number of times the user appears in different susceptible regions, a characteristic indicating the average number of times the user appears in different susceptible regions, a characteristic indicating the maximum number of times the user appears in different susceptible regions, a characteristic indicating the minimum number of times the user appears in different susceptible regions, and a characteristic indicating the standard deviation of the number of times the user appears in different susceptible regions. It should be understood that the number of times a user appears in different susceptible regions may vary.
[0084] As an example, the characteristics of the statistical values regarding the number of times a user appears in susceptible areas of different risk levels may include at least one of the following: a characteristic indicating the total number of times the user has appeared in different susceptible areas that positive sample users have cumulatively appeared m times; a characteristic indicating the average number of times the user has appeared in different susceptible areas that positive sample users have cumulatively appeared m times; a characteristic indicating the maximum number of times the user has appeared in different susceptible areas that positive sample users have cumulatively appeared m times; a characteristic indicating the minimum number of times the user has appeared in different susceptible areas that positive sample users have cumulatively appeared m times; and a characteristic indicating the standard deviation of the number of times the user has appeared in different susceptible areas that positive sample users have cumulatively appeared m times. Here, m is an integer greater than 0, and it should be understood that m can have multiple values. As an example, if a positive sample user appears in a base station area during a time period, it can be recorded as the positive sample user appearing once. The different susceptible areas that positive sample users have cumulatively appeared m times are the individual susceptible areas among all susceptible areas that positive sample users have cumulatively appeared m times.
[0085] As an example, aggregated features regarding a user and a positive sample user appearing in the same base station area during the same time period may include: features of the statistical value of the number of times the user and the positive sample user appear in the same base station area during the same time period.
[0086] As an example, the step of generating a feature for the statistical value of the number of times a user and a positive sample user appear in the same base station area within the same time period may include: constructing a dictionary of all activity trajectories of all positive sample users before their diagnosis time based on the user trajectory table and the confirmed user table; then, for each user, determining whether the user's trajectory point overlaps with the trajectory point of a positive sample user in the dictionary whose diagnosis time was before a specific time, and statistically analyzing the overlapping trajectory points to obtain a feature for the statistical value of the number of times the user and a positive sample user appear in the same base station area within the same time period. Each element in the dictionary is a trajectory point representing a positive sample user appearing in a base station area within a time period. In fact, a user's trajectory point can be defined by both the time period and the base station area. It should be understood that the specific time corresponding to each user can be obtained according to the aforementioned embodiments.
[0087] As an example, the characteristics of the statistical value regarding the number of times the user and positive sample users appear in the same base station area during the same time period may include at least one of the following: a characteristic indicating the total number of times the user and different positive sample users appear in the same base station area during the same time period; a characteristic indicating the mean of the number of times the user and different positive sample users appear in the same base station area during the same time period; a characteristic indicating the variance of the number of times the user and different positive sample users appear in the same base station area during the same time period; and a characteristic indicating the total number of positive sample users who appear in the same base station area (i.e., the same base station area) consecutively with the user for k time periods. Here, k is an integer greater than 0. It should be understood that k can be multiple values, for example, 1, 2, ..., 7.
[0088] Considering that mobile terminal base station positioning data can describe an individual's activity trajectory at the user level, which is of great significance for identifying high-risk groups who may be infected with a certain infectious disease, according to an exemplary embodiment of the present invention, based on base station positioning data and supplemented with other relevant dimensions of data, a machine learning model capable of accurately screening high-risk groups for infectious diseases is constructed by calculating features such as each user's activity trajectory, activity range, activity level, duration and distance of contact with confirmed users, whether they have been to susceptible areas, social data with confirmed users, and basic information of the user (e.g., age, gender, economic status, etc.). This model can quickly and effectively identify the most likely infected group among all mobile terminal users, so as to achieve timely detection and effective prevention of the continued spread of the disease.
[0089] Furthermore, as an example, the method for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention may further include: performing basic checks and cleaning on the data in each data table included in the training dataset before step S30.
[0090] As an example, the steps for performing basic checks and cleaning on the data in each data table may include at least one of the following processes:
[0091] Perform row cleaning for each data table. For example, for each data table, check whether the number of columns in each row is the same, whether there are duplicates, missing columns, misalignments, etc.
[0092] Perform column cleaning for each data table. For example, for each data table, check whether each column has the same variable type, and calculate the maximum, minimum, mean, and standard deviation, and calculate the fill rate for missing values.
[0093] Perform a time check on each data table. For example, for each data table, check whether the time column of the data table is within a specified time range, such as whether it is 20 consecutive days of data.
[0094] For each data table, character encoding can be performed. For example, for each data table, character data can be first standardized in terms of capitalization, and then encoded.
[0095] For each data table, missing values can be handled. For example, for each data table, missing values can be filled with the mean or 0. However, when filling the mean for time series data, care should be taken not to open the time window backward.
[0096] Figure 2 A flowchart illustrating a method for screening high-risk populations for infectious diseases according to an exemplary embodiment of the present invention is shown.
[0097] Reference Figure 2 In step S50, a prediction dataset about the users to be screened is obtained, wherein the prediction dataset includes trajectory information of the users to be screened, and the trajectory information of the users to be screened is obtained based on mobile terminal related data of the users to be screened.
[0098] It should be understood that the method of obtaining the prediction dataset of the users to be screened in step S50 is similar to the method of obtaining the training dataset in step S10, and will not be repeated here.
[0099] In step S60, features are extracted for each user to be screened based on the predicted dataset.
[0100] It should be understood that the feature extraction method in step S60 is the same as that in step S30, and will not be described again here.
[0101] In step S70, the high-risk infection population screening model for a specified type of infectious disease, generated by executing the method for generating a high-risk infection population screening model for infectious diseases as described in the exemplary embodiment above, is used to predict the risk level of the user to be screened from contracting the specified type of infectious disease based on the extracted features.
[0102] Specifically, the extracted features corresponding to each user to be screened are input into the screening model, and the risk level of the user being infected with the specified type of infectious disease is obtained from the output of the screening model.
[0103] In step S80, the predicted risk level of a user contracting the specified type of infectious disease is output. Therefore, only users with a higher risk level need to be observed and diagnosed.
[0104] As an example, the system can also output the user's ranking results in descending order of predicted risk level.
[0105] As another example, the risk level of only users whose predicted risk level meets preset conditions can be output. For example, the preset conditions could be that the risk level is higher than a preset threshold, or that the risk level is in the top N or top M% of all users to be screened, where N is an integer greater than 0 and M is a number greater than 0.
[0106] As another example, the risk level of the users to be screened that meet the preset conditions can be output only, and the sorting results can be output in descending order of the predicted risk level.
[0107] Figure 3 A block diagram of a system for generating a screening model for high-risk populations of infectious diseases according to an exemplary embodiment of the present invention is shown.
[0108] like Figure 3 As shown, the system for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention includes: a dataset acquisition device 10, a sample table establishment device 20, a feature extraction device 30, and a training device 40.
[0109] Specifically, the dataset acquisition device 10 is adapted to acquire a training dataset, wherein the training dataset includes user trajectory information, which is obtained based on user mobile terminal related data.
[0110] The sample table creation device 20 is adapted to create a sample table, wherein each sample in the sample table includes a user identifier and a sample label, the sample label indicating whether the user is a positive sample user who has been diagnosed with / suspected infection with a specified type of infectious disease or a normal negative sample user.
[0111] The feature extraction device 30 is adapted to extract features for each sample in the sample table based on the training dataset, and to incorporate the extracted features into the sample table.
[0112] The training device 40 is adapted to use machine learning algorithms to train a machine learning model based on a sample table incorporating features, thereby generating a screening model for high-risk groups of the specified type of infectious disease.
[0113] As an example, the user trajectory information may include: the base station identifier of the base station used by the user's mobile terminal in each time period, wherein each time period is obtained by dividing a specific time span according to a preset time granularity.
[0114] As an example, the base station used by a user's mobile terminal in each time period can be: the base station used by the user's mobile terminal for the longest time in each time period, or the base station used by the user's mobile terminal at a specified time point in each time period.
[0115] As an example, the dataset acquisition device 10 may be adapted to acquire signaling data of communication between a user's mobile terminal and a base station within the specific time span; and based on the signaling data, acquire the base station ID of the base station used by each user's mobile terminal in each time period of the specific time span.
[0116] As an example, the training dataset may include at least one of the following tables: a table of confirmed users, including the user IDs of positive sample users and their diagnosis times; a user trajectory table, including the user IDs and the base station IDs of the base stations used by the user's mobile terminal in each time period; a user information table, including the user IDs and the user's attribute information; a base station table, including the base station IDs and the geographical location information of the base stations; a user communication record table, including the user IDs and the communication records between the user and other users of mobile terminals using the user's mobile terminal; and a user address book information table, including the user IDs and the user IDs of contacts in the address book of at least one application of the user's mobile terminal.
[0117] As an example, the user information table may include at least one of the following user attributes: package tariff, package data usage, package call duration, package SMS message count, monthly internet data usage, monthly call duration, monthly SMS message count, monthly call charges, average monthly call duration, average monthly internet data usage, average monthly SMS message count, average monthly call charges, whether it is a group user, mobile phone number registration location, number of years of network access, age, and gender.
[0118] As an example, the feature extraction device 30 may be adapted to directly process the information in the data table included in the training dataset into basic features corresponding to each user ID; and / or, generate derived features corresponding to each user ID based on the information in the data table included in the training dataset.
[0119] As an example, the derived features may include at least one of the following: aggregated features regarding the user's activity level, aggregated features regarding the user's social intimacy with positive sample users, aggregated features regarding the user and positive sample users appearing in the same base station area during the same time period, and aggregated features regarding the user appearing in a susceptible area, wherein the base station area where the user appears during a time period is the area corresponding to the base station used by the user's mobile terminal during that time period, and the susceptible area is the base station area where each positive sample user has appeared in each time period.
[0120] As an example, aggregated features regarding a user's activity level may include at least one of the following: features indicating the number of all base station areas the user has visited before a specific time; features indicating the number of all provinces / cities the user has visited before a specific time; features indicating the maximum longitude of all base station areas the user has visited before a specific time; features indicating the minimum longitude of all base station areas the user has visited before a specific time; features indicating the maximum latitude of all base station areas the user has visited before a specific time; features indicating the minimum latitude of all base station areas the user has visited before a specific time; features indicating the user's activity distance before a specific time; and features indicating whether the user is from another province.
[0121] As an example, if the user is a positive sample user, the specific time corresponding to them can be the time of their diagnosis; if the user is a negative sample user, the specific time corresponding to them can be the time assigned to them according to a specific rule. Further, as an example, the specific time assigned to the user according to the specific rule can be the last day of the specific time span, or the specific time corresponding to each of the negative sample users can be uniformly set according to the distribution of the diagnosis times of all positive sample users.
[0122] As an example, aggregated features regarding the social intimacy between a user and positive sample users may include: features regarding the statistical value of the number of positive sample users with different social intimacy with that user, wherein the social intimacy between users is determined based on call records between users and / or the user's address book information.
[0123] As an example, aggregated features regarding a user and a positive sample user appearing in the same base station area during the same time period may include: features of the statistical value of the number of times the user and the positive sample user appear in the same base station area during the same time period.
[0124] As an example, the aggregated features of a user's presence in a vulnerable area may include at least one of the following: a feature of the statistical value of the number of times the user appears in each vulnerable area, and a feature of the statistical value of the number of times the user appears in vulnerable areas of different risk levels, wherein the risk level of the vulnerable area is related to the number of times the positive sample user has appeared.
[0125] As an example, the feature extraction device 30 may be adapted to construct a dictionary of all activity trajectories of all positive sample users before their diagnosis time based on the user trajectory table and the confirmed user table; for each user, it determines whether the user's trajectory points overlap with the trajectory points of positive sample users in the dictionary whose diagnosis time was before a specific time, and performs statistics on the overlapping trajectory points to obtain a feature of the statistical value of the number of times the user and positive sample users appeared in the same base station area in the same time period, wherein each element in the dictionary is a trajectory point used to characterize a positive sample user appearing in a base station area in a time period.
[0126] As an example, the machine learning algorithm may be a model fusion algorithm.
[0127] It should be understood that the specific implementation of the system for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention can be referred to in conjunction with... Figure 1 The specific implementation methods described will not be elaborated here.
[0128] Figure 4 A block diagram of a system for screening high-risk populations for infectious diseases according to an exemplary embodiment of the present invention is shown.
[0129] like Figure 4 As shown, a system for screening high-risk groups for infectious diseases according to an exemplary embodiment of the present invention may include: a dataset acquisition device 50, a feature extraction device 60, a prediction device 70, and an output device 80.
[0130] Specifically, the dataset acquisition device 50 is adapted to acquire a predictive dataset about the users to be screened, wherein the predictive dataset includes trajectory information of the users to be screened, and the trajectory information of the users to be screened is obtained based on mobile terminal-related data of the users to be screened.
[0131] The feature extraction device 60 is adapted to extract features for each user to be screened based on the prediction dataset.
[0132] The prediction device 70 is adapted to use the high-risk infection population screening model for a specified type of infectious disease generated by the system for generating high-risk infection population screening models for infectious diseases as described above, and to predict the risk level of the user to be screened for infection with the specified type of infectious disease based on the extracted features.
[0133] The output device 80 is adapted to output the predicted risk level of a user contracting the specified type of infectious disease.
[0134] As an example, the output device 80 may be adapted to output the user's ranking results in descending order of predicted risk levels.
[0135] As another example, the output device 80 may be adapted to output only the risk level of users whose predicted risk level meets preset conditions.
[0136] It should be understood that the specific implementation of the system for screening high-risk groups for infectious diseases according to exemplary embodiments of the present invention can be referred to in conjunction with... Figure 2 The specific implementation methods described will not be elaborated here.
[0137] The system for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention includes devices that can be configured as software, hardware, firmware, or any combination thereof to perform specific functions. For example, these devices may correspond to dedicated integrated circuits, pure software code, or modules combining software and hardware. Furthermore, one or more functions implemented by these devices may also be uniformly executed by components in a physical entity device (e.g., a processor, client, or server).
[0138] The devices included in the system for screening high-risk populations for infectious diseases according to exemplary embodiments of the present invention can be configured as software, hardware, firmware, or any combination thereof to perform specific functions. For example, these devices may correspond to dedicated integrated circuits, pure software code, or modules combining software and hardware. Furthermore, one or more functions implemented by these devices may also be uniformly executed by components in a physical entity device (e.g., a processor, client, or server).
[0139] It should be understood that the method for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention can be implemented by a program recorded on a computer-readable medium. For example, according to an exemplary embodiment of the present invention, a computer-readable medium for generating a screening model for high-risk groups of infectious diseases can be provided, wherein a computer program for performing the following method steps is recorded on the computer-readable medium: acquiring a training dataset, wherein the training dataset includes user trajectory information, wherein the user trajectory information is obtained based on user mobile terminal related data; establishing a sample table, wherein each sample in the sample table includes a user identifier and a sample label, the sample label indicating whether the user is a positive sample user diagnosed with / suspected infection with a specified type of infectious disease or a normal negative sample user; extracting features for each sample in the sample table based on the training dataset, and incorporating the extracted features into the sample table; using a machine learning algorithm, training a machine learning model based on the sample table with the incorporated features, to generate a screening model for high-risk groups of the specified type of infectious disease.
[0140] It should be understood that the method for screening high-risk groups for infectious diseases according to exemplary embodiments of the present invention can be implemented by a program recorded on a computer-readable medium. For example, according to exemplary embodiments of the present invention, a computer-readable medium for screening high-risk groups for infectious diseases can be provided, wherein a computer program for performing the following method steps is recorded on the computer-readable medium: acquiring a prediction dataset about users to be screened, wherein the prediction dataset includes trajectory information of the users to be screened, wherein the trajectory information of the users to be screened is obtained based on mobile terminal-related data of the users to be screened; extracting features for each user to be screened based on the prediction dataset; predicting the risk level of the user to be screened for the specified type of infectious disease based on the extracted features using a high-risk group screening model for a specified type of infectious disease generated by performing the method as described in the exemplary embodiments above; and outputting the predicted risk level of the user for the specified type of infectious disease.
[0141] The computer program in the aforementioned computer-readable medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, and servers. It should be noted that the computer program can also be used to perform additional steps besides those described above, or to perform more specific processing while performing the above steps. The content of these additional steps and further processing has been described in reference to... Figure 1 and Figure 2 The above has already been described, and will not be repeated here to avoid repetition.
[0142] It should be noted that the system for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention can rely entirely on the operation of a computer program to realize the corresponding functions. That is, each device corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a special software package (e.g., a lib library) to realize the corresponding functions.
[0143] On the other hand, the various devices included in the system for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention can also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segment for performing the corresponding operation can be stored in a computer-readable medium such as a storage medium, so that a processor can perform the corresponding operation by reading and running the corresponding program code or code segment.
[0144] It should be noted that the system for screening high-risk groups for infectious diseases according to an exemplary embodiment of the present invention can rely entirely on the operation of a computer program to achieve the corresponding functions. That is, each device corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a special software package (e.g., a lib library) to achieve the corresponding functions.
[0145] On the other hand, the various devices included in the system for screening high-risk groups for infectious diseases according to exemplary embodiments of the present invention can also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segment for performing the corresponding operation can be stored in a computer-readable medium such as a storage medium, so that a processor can perform the corresponding operation by reading and running the corresponding program code or code segment.
[0146] For example, an exemplary embodiment of the present invention can also be implemented as a computing device, which includes a storage component and a processor. The storage component stores a set of computer-executable instructions, which, when executed by the processor, execute a method for generating a screening model for high-risk groups of infectious diseases and / or a method for screening high-risk groups of infectious diseases.
[0147] Specifically, the computing device can be deployed on a server or client, or on a node device in a distributed network environment. Furthermore, the computing device can be a PC, tablet, personal digital assistant, smartphone, web application, or other device capable of executing the aforementioned set of instructions.
[0148] Here, the computing device is not necessarily a single computing device, but can be any collection of devices or circuits capable of executing the above instructions (or instruction sets) individually or in combination. The computing device can also be part of an integrated control system or system manager, or can be configured to interconnect with a portable electronic device locally or remotely (e.g., via wireless transmission) through an interface.
[0149] In the computing device, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, etc.
[0150] Some operations described in the method for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention can be implemented by software, some operations can be implemented by hardware, and these operations can also be implemented by a combination of software and hardware.
[0151] Some operations described in the method for screening high-risk groups for infectious diseases according to exemplary embodiments of the present invention can be implemented by software, some operations can be implemented by hardware, and these operations can also be implemented by a combination of software and hardware.
[0152] The processor can execute instructions or code stored in one of its storage components, which can also store data. The instructions and data can also be sent and received over a network via a network interface device, which can employ any known transport protocol.
[0153] Storage components can be integrated with the processor, for example, by placing RAM or flash memory within an integrated circuit microprocessor. Alternatively, storage components can include separate devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. Storage components and the processor can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., enabling the processor to read files stored in the storage component.
[0154] In addition, the computing device may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the computing device may be interconnected via a bus and / or network.
[0155] The operations involved in the method for generating a screening model for high-risk populations of infectious diseases according to an exemplary embodiment of the present invention can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally integrated into a single logical device or operate according to non-precise boundaries.
[0156] For example, as described above, a computing device for generating a screening model for high-risk groups of infectious diseases according to an exemplary embodiment of the present invention may include a storage component and a processor, wherein the storage component stores a set of computer-executable instructions, and when the set of computer-executable instructions is executed by the processor, the following steps are performed: acquiring a training dataset, wherein the training dataset includes user trajectory information, wherein the user trajectory information is obtained based on user mobile terminal related data; establishing a sample table, wherein each sample in the sample table includes a user identifier and a sample label, the sample label indicating whether the user is a positive sample user who has been diagnosed with / suspected infection with a specified type of infectious disease or a normal negative sample user; extracting features for each sample in the sample table based on the training dataset, and incorporating the extracted features into the sample table; using a machine learning algorithm, training a machine learning model based on the sample table with the incorporated features, to generate a screening model for high-risk groups of the specified type of infectious disease.
[0157] The operations involved in the method for screening high-risk populations for infectious diseases according to exemplary embodiments of the present invention can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally integrated into a single logical device or operate according to non-precise boundaries.
[0158] For example, as described above, a computing device for screening high-risk groups for infectious diseases according to an exemplary embodiment of the present invention may include a storage component and a processor, wherein the storage component stores a set of computer-executable instructions, and when the set of computer-executable instructions is executed by the processor, the following steps are performed: acquiring a prediction dataset about users to be screened, wherein the prediction dataset includes trajectory information of users to be screened, wherein the trajectory information of users to be screened is obtained based on mobile terminal-related data of users to be screened; extracting features for each user to be screened based on the prediction dataset; using a screening model for high-risk groups for a specified type of infectious disease generated by executing the method described in the exemplary embodiment above, predicting the risk level of the user to be screened from contracting the specified type of infectious disease based on the extracted features; and outputting the predicted risk level of the user from contracting the specified type of infectious disease.
[0159] The foregoing has described various exemplary embodiments of the present invention. It should be understood that the above description is merely exemplary and not exhaustive, and the present invention is not limited to the disclosed exemplary embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for generating a screening model for high-risk populations of infectious diseases, wherein, The method includes: Obtain a training dataset, wherein the training dataset includes user trajectory information, wherein the user trajectory information is obtained based on user mobile terminal related data, and the user trajectory information includes: the base station identifier of the base station used by the user's mobile terminal in each time period; Establish a sample table, wherein each sample in the sample table includes a user identifier and a sample label, and the sample label indicates whether the user is a positive sample user who has been diagnosed with / suspected infection with a specified type of infectious disease or a normal negative sample user; Based on the training dataset, features are extracted for each sample in the sample table, and the extracted features are incorporated into the sample table. The sample table is obtained by concatenating the extracted derived features to the corresponding samples to obtain a sample table with incorporated features. The derived features include at least one of the following: aggregated features about the user's activity level, aggregated features about the social intimacy between the user and positive sample users, aggregated features about the user and positive sample users appearing in the same base station area at the same time, and aggregated features about the user appearing in a susceptible area. Using machine learning algorithms, a machine learning model is trained based on a sample table incorporating features to generate a screening model for high-risk groups of the specified type of infectious disease.
2. The method as described in claim 1, wherein, Each time period is obtained by dividing a specific time span according to a preset time granularity.
3. The method as described in claim 2, wherein, The base stations used by a user's mobile terminal in each time period are: the base station used by the user's mobile terminal for the longest time in each time period, or the base station used by the user's mobile terminal at a specified time point in each time period.
4. The method of claim 2, wherein, The training dataset includes at least one item from the following data tables: The table of confirmed users includes the user IDs of positive sample users and their diagnosis times; The user trajectory table includes the user ID and the base station ID of the base station used by the user's mobile terminal in each time period; The user information table includes the user ID and the user's attribute information; The base station table includes the base station ID and the geographical location information of the base station; The user communication record table includes the user ID and the communication records between the user and other users using mobile terminals; The user address book information table includes the user ID and the user IDs of contacts in the address book of at least one application on the user's mobile terminal.
5. The method of claim 4, wherein, The steps for extracting features for each sample in the sample table based on the training dataset include: The information in the data tables included in the training dataset is directly processed into basic features corresponding to each user ID; And / or, based on the information in the data tables included in the training dataset, generate derived features corresponding to each user ID. Among them, the base station area where a user appears in a certain time period is the area corresponding to the base station used by the user's mobile terminal in that time period, and the susceptible area is the base station area where each positive sample user appears in each time period.
6. The method of claim 5, wherein, The aggregated features regarding a user's activity level include at least one of the following: features indicating the number of all base station areas the user has visited before a specific time; features indicating the number of all provinces / cities the user has visited before a specific time; features indicating the maximum longitude of all base station areas the user has visited before a specific time; features indicating the minimum longitude of all base station areas the user has visited before a specific time; features indicating the maximum latitude of all base station areas the user has visited before a specific time; features indicating the minimum latitude of all base station areas the user has visited before a specific time; features indicating the user's activity distance before a specific time; and features indicating whether the user is from another province. The aggregated features regarding the social intimacy between a user and positive sample users include: a feature of the statistical value of the number of positive sample users with different social intimacy with the user, wherein the social intimacy between users is determined based on the call records between users and / or the user's address book information; The aggregated features regarding the occurrence of a user and a positive sample user in the same base station area within the same time period include: the statistical value of the number of times the user and the positive sample user appeared in the same base station area within the same time period; Aggregated features regarding a user's presence in susceptible areas include at least one of the following: a feature representing the statistical value of the number of times the user appears in each susceptible area; a feature representing the statistical value of the number of times the user appears in susceptible areas of different risk levels. The level of danger in susceptible areas is related to the number of times positive sample users have been present. If the user is a positive sample user, the specific time corresponding to that user is the time of diagnosis; if the user is a negative sample user, the specific time corresponding to that user is the time assigned to that user according to specific rules.
7. The method of claim 6, wherein, The steps for generating features that show the number of times a user and a positive sample user appear in the same base station area within the same time period include: Based on the user trajectory table and the confirmed user table, a dictionary is constructed for all activity trajectories of all positive sample users before their diagnosis time, wherein each element in the dictionary is a trajectory point representing a positive sample user appearing in a base station area within a time period; For each user, it is determined whether the user's trajectory points overlap with the trajectory points of positive sample users whose diagnosis time was before a specific time in the dictionary, and the overlapping trajectory points are statistically analyzed to obtain the statistical value of the number of times the user and positive sample users appeared in the same base station area in the same time period.
8. The method of claim 4, wherein, The user information table includes at least one of the following user attribute information: Package rates, data allowance, call duration, SMS message count, monthly data usage, monthly call duration, monthly SMS message count, monthly call charges, average monthly call duration, average monthly data usage, average monthly SMS message count, average monthly call charges, whether it is a corporate user, mobile number registration location, number of years with the network, age, and gender.
9. The method of claim 2, wherein, The steps to obtain the training dataset include: Obtain signaling data of communication between the user's mobile terminal and the base station within the specified time span; Based on the signaling data, obtain the base station ID of the base station used by each user's mobile terminal in each time period of the specific time span.
10. The method of claim 1, wherein, The machine learning algorithm mentioned is a model fusion algorithm.
11. The method of claim 6, wherein, The user's specific time is assigned to the last day of the specific time span according to the specific rules, or the specific time corresponding to each of the negative sample users is uniformly set according to the distribution of the diagnosis time of all positive sample users.
12. A method for screening high-risk populations for infectious diseases, wherein, The method includes: Obtain a predictive dataset about users to be screened, wherein the predictive dataset includes trajectory information of the users to be screened, and the trajectory information of the users to be screened is obtained based on mobile terminal-related data of the users to be screened; Based on the predicted dataset, features are extracted for each user to be screened. Using a high-risk infection population screening model for a specified type of infectious disease generated by performing the method as described in any one of claims 1 to 11, the risk of a user to be screened contracting the specified type of infectious disease is predicted based on extracted features; Output the predicted risk level of a user contracting the specified type of infectious disease.
13. The method of claim 12, wherein, The steps for outputting the predicted risk level of a user contracting the specified type of infectious disease include: Output the users' ranking results in descending order of predicted risk level; And / or, output only the risk level of users whose predicted risk level meets preset conditions.
14. A system for generating a screening model for high-risk populations of infectious diseases, wherein, The system includes: A dataset acquisition device is adapted to acquire a training dataset, wherein the training dataset includes user trajectory information, wherein the user trajectory information is obtained based on data related to the user's mobile terminal, and the user trajectory information includes: the base station identifier of the base station used by the user's mobile terminal in each time period; A sample table creation device is suitable for creating a sample table, wherein each sample in the sample table includes a user identifier and a sample label, and the sample label indicates whether the user is a positive sample user who has been diagnosed with / suspected infection with a specified type of infectious disease or a normal negative sample user. A feature extraction device is adapted to extract features for each sample in the sample table based on the training dataset, and to incorporate the extracted features into the sample table. The sample table is obtained by concatenating the extracted derived features to the corresponding samples to obtain a sample table incorporating the features. The derived features include at least one of the following: aggregated features about the user's activity level, aggregated features about the social intimacy between the user and positive sample users, aggregated features about the user and positive sample users appearing in the same base station area at the same time, and aggregated features about the user appearing in a susceptible area. The training device is adapted to use machine learning algorithms to train a machine learning model based on a sample table incorporating features, thereby generating a screening model for high-risk groups of the specified type of infectious disease.
15. The system of claim 14, wherein, Each time period is obtained by dividing a specific time span according to a preset time granularity.
16. The system of claim 15, wherein, The base stations used by a user's mobile terminal in each time period are: the base station used by the user's mobile terminal for the longest time in each time period, or the base station used by the user's mobile terminal at a specified time point in each time period.
17. The system of claim 15, wherein, The training dataset includes at least one item from the following data tables: The table of confirmed users includes the user IDs of positive sample users and their diagnosis times; The user trajectory table includes the user ID and the base station ID of the base station used by the user's mobile terminal in each time period; The user information table includes the user ID and the user's attribute information; The base station table includes the base station ID and the geographical location information of the base station; The user communication record table includes the user ID and the communication records between the user and other users using mobile terminals; The user address book information table includes the user ID and the user IDs of contacts in the address book of at least one application on the user's mobile terminal.
18. The system of claim 17, wherein, The feature extraction device is adapted to directly process the information in the data tables included in the training dataset into basic features corresponding to each user ID; And / or, based on the information in the data tables included in the training dataset, generate derived features corresponding to each user ID. Among them, the base station area where a user appears in a certain time period is the area corresponding to the base station used by the user's mobile terminal in that time period, and the susceptible area is the base station area where each positive sample user appears in each time period.
19. The system of claim 18, wherein, The aggregated features regarding a user's activity level include at least one of the following: features indicating the number of all base station areas the user has visited before a specific time; features indicating the number of all provinces / cities the user has visited before a specific time; features indicating the maximum longitude of all base station areas the user has visited before a specific time; features indicating the minimum longitude of all base station areas the user has visited before a specific time; features indicating the maximum latitude of all base station areas the user has visited before a specific time; features indicating the minimum latitude of all base station areas the user has visited before a specific time; features indicating the user's activity distance before a specific time; and features indicating whether the user is from another province. The aggregated features regarding the social intimacy between a user and positive sample users include: a feature of the statistical value of the number of positive sample users with different social intimacy with the user, wherein the social intimacy between users is determined based on the call records between users and / or the user's address book information; The aggregated features regarding the occurrence of a user and a positive sample user in the same base station area within the same time period include: the statistical value of the number of times the user and the positive sample user appeared in the same base station area within the same time period; Aggregated features regarding a user's presence in susceptible areas include at least one of the following: a feature representing the statistical value of the number of times the user appears in each susceptible area; a feature representing the statistical value of the number of times the user appears in susceptible areas of different risk levels. The level of danger in susceptible areas is related to the number of times positive sample users have been present. If the user is a positive sample user, the specific time corresponding to that user is the time of diagnosis; if the user is a negative sample user, the specific time corresponding to that user is the time assigned to that user according to specific rules.
20. The system of claim 19, wherein, The feature extraction device is adapted to construct a dictionary of all activity trajectories of all positive sample users prior to their diagnosis time based on the user trajectory table and the confirmed user table; for each user, it determines whether the user's trajectory points overlap with the trajectory points of positive sample users in the dictionary whose diagnosis time was prior to a specific time, and statistically analyzes the overlapping trajectory points to obtain a feature of the statistical value of the number of times the user and positive sample users appeared in the same base station area during the same time period. Each element in the dictionary is a trajectory point representing a positive sample user appearing in a base station area within a time period.
21. The system of claim 17, wherein, The user information table includes at least one of the following user attribute information: Package rates, data allowance, call duration, SMS message count, monthly data usage, monthly call duration, monthly SMS message count, monthly call charges, average monthly call duration, average monthly data usage, average monthly SMS message count, average monthly call charges, whether it is a corporate user, mobile number registration location, number of years with the network, age, and gender.
22. The system of claim 15, wherein, The dataset acquisition device is adapted to acquire signaling data of communication between a user's mobile terminal and a base station within the specific time span; and based on the signaling data, acquire the base station ID of the base station used by each user's mobile terminal in each time period of the specific time span.
23. The system of claim 14, wherein, The machine learning algorithm mentioned is a model fusion algorithm.
24. The system of claim 19, wherein, The user's specific time is assigned to the last day of the specific time span according to the specific rules, or the specific time corresponding to each of the negative sample users is uniformly set according to the distribution of the diagnosis time of all positive sample users.
25. A system for screening high-risk populations for infectious diseases, wherein, The system includes: A dataset acquisition device is adapted to acquire a predictive dataset about users to be screened, wherein the predictive dataset includes trajectory information of the users to be screened, wherein the trajectory information of the users to be screened is obtained based on mobile terminal-related data of the users to be screened; The feature extraction device is adapted to extract features for each user to be screened based on the prediction dataset; A prediction device is adapted to predict the risk level of a user to be screened for the specified type of infectious disease based on extracted features, using a screening model for a high-risk population for a specified type of infectious disease generated by the system as described in any one of claims 14 to 24. An output device adapted to output the predicted risk level of a user contracting the specified type of infectious disease.
26. The system of claim 25, wherein, The output device is adapted to output the user's ranking results in descending order of predicted risk level; And / or, the output device is adapted to output only the risk level of users whose predicted risk level meets preset conditions.
27. A system comprising at least one computing device and at least one storage device for storing instructions, wherein, When the instructions are executed by the at least one computing device, the at least one computing device performs the method for generating a screening model for high-risk groups of infectious diseases as described in any one of claims 1 to 11 and / or the method for screening high-risk groups of infectious diseases as described in any one of claims 12 to 13.
28. A computer-readable storage medium for storing instructions, wherein, When the instructions are executed by at least one computing device, the at least one computing device is prompted to perform the method for generating a screening model for high-risk groups of infectious diseases as described in any one of claims 1 to 11 and / or the method for screening high-risk groups of infectious diseases as described in any one of claims 12 to 13.
Citation Information
Patent Citations
Method for tracking infection sources and predicting trends of infectious diseases by utilizing mobile phone tracks
CN105740615A
Risk model training method and apparatus, risk identification method and apparatus, equipment, and medium
CN108520343A
Disease prediction method, device, medium and electronic equipment
CN108986921A
Risk prediction model training method and device, risk prediction method and device, medium and equipment
CN110674979A