A data processing method, device, apparatus, and storage medium

By acquiring statistical parameters and scoring methods from population mobility models, the problem of the inability to effectively analyze various models in existing technologies has been solved, enabling accurate evaluation and selection in different scenarios.

CN116796115BActive Publication Date: 2026-05-19CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
Filing Date
2022-09-01
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, there are many studies on population flow models for specific scenarios, but there is no way to effectively analyze and evaluate the statistical results of various population flow models.

Method used

A data processing method is provided that by obtaining statistical parameters of a statistical model, including location information, the statistical time period of the origin and destination, the target scenario is determined, and the model is scored based on this information to evaluate its accuracy.

Benefits of technology

It enables accurate scoring of statistical models, improves the accuracy of evaluation, and allows for the selection of appropriate models for population flow statistics and prediction in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116796115B_ABST
    Figure CN116796115B_ABST
Patent Text Reader

Abstract

The application discloses a data processing method, device and equipment and a storage medium. The method comprises the following steps: obtaining a statistical parameter of a first statistical model; the statistical parameter comprises at least the following: place information, a statistical time period of a starting place, and a statistical time period of a destination; determining a target scene corresponding to the first statistical model based on at least the place information; and determining a score of the first statistical model based on at least the statistical time period of the starting place, the statistical time period of the destination and the target scene; and the score of the first statistical model is used for evaluating the first statistical model. According to the scheme, the statistical model can be evaluated, and the evaluation accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, including but not limited to data processing methods, apparatus, devices, and storage media. Background Technology

[0002] With the development of urban transportation and the market economy, population mobility has become very frequent, and research on population mobility phenomena has gradually increased, including exploring the purpose of population mobility, conducting population migration statistics in special scenarios in a certain way, and predicting population mobility trends.

[0003] In related technologies, there are many studies on population flow models for specific scenarios, but in the face of numerous population flow models, it is impossible to analyze the statistical results of various population flow models. Summary of the Invention

[0004] This application provides a data processing method, apparatus, device, and storage medium that can evaluate statistical models with high accuracy.

[0005] The technical solution of this application is implemented as follows:

[0006] This application provides a data processing method, the method comprising:

[0007] Obtain the statistical parameters of the first statistical model; the statistical parameters include at least: location information, the statistical time period of the origin, and the statistical time period of the destination;

[0008] Based at least on the location information, the target scene corresponding to the first statistical model is determined;

[0009] The score of the first statistical model is determined based at least on the statistical period of the origin, the statistical period of the destination, and the target scenario; the score of the first statistical model is used to evaluate the first statistical model.

[0010] This application provides a data processing apparatus, the apparatus comprising:

[0011] The obtaining unit is used to obtain statistical parameters of the first statistical model; the statistical parameters include at least: location information, statistical time period of the origin, and statistical time period of the destination;

[0012] The first determining unit is configured to determine the target scene corresponding to the first statistical model based at least on the location information.

[0013] The second determining unit is used to determine the score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scenario; the score of the first statistical model is used to evaluate the first statistical model.

[0014] This application also provides an electronic device, including: a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the above-described data processing method.

[0015] This application also provides a storage medium on which a computer program is stored, which, when executed by a processor, implements the above-described data processing method.

[0016] The data processing method, apparatus, device, and storage medium provided in this application include: obtaining statistical parameters of a first statistical model; the statistical parameters include at least: location information, a statistical time period of the origin, and a statistical time period of the destination; determining a target scene corresponding to the first statistical model based at least on the location information; determining a score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scene; the score of the first statistical model is used to evaluate the first statistical model. For the solution of this application, the statistical parameters of the first statistical model are first obtained, and then the target scene corresponding to the first statistical model is determined based on the statistical parameters; then the score of the first statistical model is determined based on the statistical parameters and the target scene. It can be seen that the data processing method of this application can score the statistical model, and because the scene information of the statistical model is combined in the process of determining the score, the accuracy of the evaluation is also relatively high. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of an optional structure of the data processing system provided in an embodiment of this application;

[0018] Figure 2 An optional flowchart illustrating the data processing method provided in the embodiments of this application.

[0019] Figure 3 A schematic flowchart of an optional data processing method provided in an embodiment of this application;

[0020] Figure 4 A schematic flowchart of an optional data processing method provided in an embodiment of this application;

[0021] Figure 5 An optional structural diagram illustrating the relationship between parameters in the population flow model provided in this application embodiment;

[0022] Figure 6 An optional structural diagram illustrating the relationship between some parameters in the population flow model provided in this application embodiment;

[0023] Figure 7This is an optional structural diagram of the data processing procedure provided in an embodiment of this application;

[0024] Figure 8 This is a schematic diagram of an optional structure of the data processing apparatus provided in an embodiment of this application;

[0025] Figure 9 This is a schematic diagram of an optional structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of the application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0027] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0028] In the following description, the terms "first," "second," and "third" are used only to distinguish different objects and do not represent a specific order of objects, nor are they constituting a chronological order. It is understood that "first," "second," and "third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0030] This application provides data processing methods, apparatus, devices, and storage media. In practical applications, the data processing method can be implemented by a data processing apparatus, and the functional entities in the data processing apparatus can be collaboratively implemented by the hardware resources of electronic devices, such as computing resources like processors and communication resources (such as those used to support various communication methods like optical fiber and cellular networks).

[0031] The data processing method provided in this application embodiment is applied to a data processing system, which includes a first device.

[0032] The first device is used to perform: obtaining statistical parameters of a first statistical model; the statistical parameters include at least: location information, statistical time period of the origin, and statistical time period of the destination; determining the target scenario corresponding to the first statistical model based at least on the location information; determining the score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scenario; the score of the first statistical model is used to evaluate the first statistical model.

[0033] Optionally, the data processing system may also include a second device. The second device is used to perform the statistical functions of the first statistical model, obtain statistical parameters during the statistical process, and send the obtained statistical parameters to the first device.

[0034] It is understood that the statistical process of the first statistical model and the process of obtaining statistical parameters based on the statistical process can be implemented on the first device side or on the second device side; the embodiments of this application do not limit this to a single one.

[0035] It should be noted that the first device and the second device can be integrated into the same electronic device, or they can be deployed independently on different electronic devices.

[0036] As an example, the structure of a data processing system can be as follows: Figure 1 As shown, it includes: a first device 10 and a second device 20. Data can be transmitted between the first device 10 and the second device 20.

[0037] Here, the first device 10 is used to perform: obtaining statistical parameters of a first statistical model; the statistical parameters include at least: location information, statistical time period of the origin, and statistical time period of the destination; determining the target scene corresponding to the first statistical model based at least on the location information; determining the score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scene; the score of the first statistical model is used to evaluate the first statistical model.

[0038] The first device 10 can be a mobile terminal device (such as a mobile phone, tablet computer, etc.) or a non-mobile terminal device (such as a desktop computer, server, etc.) or other electronic devices with relevant data processing capabilities.

[0039] The second electronic device 20 is used to perform the statistical function of the first statistical model, obtain statistical parameters during the statistical process, and send the obtained statistical parameters to the first device 10.

[0040] The second electronic device 20 may include: mobile terminal devices (such as mobile phones, tablets, etc.) or non-mobile terminal devices (such as desktop computers, servers, etc.) and other electronic devices with relevant data processing capabilities.

[0041] Below, in conjunction with Figure 1 The schematic diagram of the data processing system shown illustrates various embodiments of the data processing methods, apparatus, devices, and storage media provided in this application.

[0042] To facilitate understanding, some technical terms will be explained first.

[0043] Statistical models are used to identify migration objects that meet the requirements.

[0044] The term "migration object" refers to anything that possesses the ability to migrate actively or passively. Examples include, but are not limited to, people and animals. It should be noted that the migration object in this application's embodiments can refer to people in general, or specifically to a particular type of person, such as university students.

[0045] For example, statistical model A is used to count the number of people who moved from Suzhou to Beijing, Shanghai, and Nanjing from March 2020 to June 2020. Specifically, statistical model A counts the number of people residing in Suzhou from January 2020 to February 2020, and counts the number of people who moved from Suzhou to Beijing, from Suzhou to Shanghai, and from Suzhou to Nanjing from March 2020 to June 2020.

[0046] In a first aspect, embodiments of this application provide a data processing method, which is applied to a data processing apparatus; wherein the data processing apparatus can be deployed in... Figure 1 The first device 10 in the process. The data processing procedure provided in the embodiments of this application will now be described using an electronic device as the executing entity.

[0047] Figure 2 This diagram illustrates a flow chart of one possible data processing method. (Refer to...) Figure 2 The data processing method shown may include, but is not limited to, the following: Figure 2 S201 to S203 are shown.

[0048] S201, The electronic device obtains the statistical parameters of the first statistical model; the statistical parameters include at least: location information, the statistical time period of the origin, and the statistical time period of the destination.

[0049] The first statistical model is used to perform statistical analysis on objects that meet the requirements based on statistical parameters. The first statistical model can be any statistical model. This application does not limit the statistical method of the first statistical model or the specific statistical requirements; these can be determined according to actual needs.

[0050] Example 1: The first statistical model can be used to count the number of people who moved from Suzhou to Beijing, Shanghai and Nanjing between March 2020 and June 2020.

[0051] Statistical parameters refer to the relevant parameters configured in the first statistical model. This application does not impose a unique limitation on the specific types of statistical parameters; they can be determined based on actual circumstances.

[0052] In one possible implementation, the statistical parameters may include: location information, the statistical time period of the origin, and the statistical time period of the destination.

[0053] Location information is used to represent the origin and destination before and after the migration. This location information may include at least one place of revelation and at least one destination. Based on Example 1, the origin is Suzhou, and the destinations include Beijing, Shanghai, and Nanjing.

[0054] The statistical period of the origin is used to characterize the time period during which statistics are collected at the origin. Based on Example 1, the statistical period of the origin can be from January 2020 to February 2020.

[0055] The statistical period for a destination is used to characterize the time period during which statistics are collected at the destination. Based on Example 1, the statistical period for a destination could be from March 2020 to June 2020.

[0056] In another possible implementation, the statistical parameters may also include: the statistical standard for the length of stay at the origin and the statistical standard for the length of stay at the destination.

[0057] Among them, the standard for the length of stay at the place of origin is used to limit the standard for the length of stay when counting at the place of origin.

[0058] For example, if the statistical standard for the length of stay at the place of origin is 3 days, then a sample (citizen A) will only be included in the statistics if the length of stay at the place of origin is greater than or equal to 3 days, that is, it is considered that citizen A stayed at the place of origin.

[0059] Otherwise, it is assumed that the citizen was merely passing through and did not actually stay at the starting point. For example, if citizen B stayed at the starting point for 30 minutes, it is assumed that citizen B was merely passing through and did not stay.

[0060] The standard for calculating the length of stay at a destination is used to define the standard for the length of stay when calculating the length of stay at the destination.

[0061] The specific details of the statistical standards for the length of stay at the destination are similar to those for the length of stay at the origin, and will not be repeated here. For specific implementation details, please refer to the detailed description of the statistical standards for the length of stay at the origin.

[0062] In another possible implementation, the statistical parameters may also include: a first distance, a first speed, and a second speed.

[0063] The first distance refers to the distance between the starting point and the destination. This application does not limit the actual length of the first distance in its embodiments; it can be configured according to actual needs.

[0064] The first velocity refers to the minimum velocity during the migration process.

[0065] The second velocity refers to the maximum velocity during the migration process.

[0066] The embodiments of this application do not limit the specific values ​​of the first speed and the second speed, which can be determined according to the actual situation. For example, they can be determined according to the type of transportation being moved.

[0067] In one possible implementation, S201 can be implemented as follows: the electronic device determines the statistical parameters of the first statistical model based on the statistical requirements of the first statistical model (location information, statistical time period of the origin, statistical time period of the destination, etc.).

[0068] In another possible implementation, other devices or electronic devices determine the statistical parameters of the first statistical model based on the statistical requirements of the first statistical model (location information, statistical time period of origin, statistical time period of destination, etc.); and send the statistical parameters to the electronic device; S201 can be implemented as follows: the electronic device receives the statistical parameters of the first statistical model sent by other devices, and obtains the location information, statistical time period of origin, and statistical time period of destination from the statistical parameters.

[0069] S202. The electronic device determines the target scene corresponding to the first statistical model based at least on the location information.

[0070] The embodiments of this application do not impose a unique limitation on the number and type of scenarios involved, and can be configured according to actual conditions.

[0071] For example, the scenarios involved in the embodiments of this application may include, but are not limited to: migration from origin, migration to destination, and flow between origin and destination.

[0072] The target scenario is the scenario corresponding to the first statistical model. This application embodiment does not limit the specific process of determining the target scenario, and it can be determined according to the actual situation.

[0073] In one possible implementation, S202 can be implemented as follows: the electronic device acquires location information from the statistical information, and determines the target scene corresponding to the first statistical model based on the origin and destination in the location information.

[0074] In another possible implementation, S202 can be implemented as follows: the electronic device acquires location information and other information (such as speed information) from the statistical information, and determines the target scene corresponding to the first statistical model based on the origin and destination and speed information in the location information.

[0075] S203. The electronic device determines the score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scenario.

[0076] The score of the first statistical model is used to evaluate the first statistical model.

[0077] For example, the score of the first statistical model is used to evaluate the completeness of the first statistical model in the statistical process; for example, the score of the first statistical model can also be used to evaluate the misclassification rate of the first statistical model in the statistical process; for example, the score of the first statistical model can also be used to evaluate the statistical cost of the first statistical model in the statistical process.

[0078] In one possible implementation, S203 can be implemented as follows: the electronic device determines the score of the first statistical model based on the statistical time period of the origin, the statistical time period of the destination, and the target scenario.

[0079] In one possible implementation, S203 can be implemented as follows: the electronic device determines the score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, the statistical standard of the dwell time of the origin, the statistical standard of the dwell time of the destination, and the target scenario.

[0080] The embodiments of this application do not specifically limit the process by which the electronic device determines the score of the first statistical model, and can be configured according to the actual situation.

[0081] The data processing scheme provided in this application includes: obtaining statistical parameters of a first statistical model; the statistical parameters include at least: location information, a statistical time period of the origin, and a statistical time period of the destination; determining a target scene corresponding to the first statistical model based at least on the location information; determining a score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scene; the score of the first statistical model is used to evaluate the first statistical model. For the scheme of this application, the statistical parameters of the first statistical model are first obtained, then the target scene corresponding to the first statistical model is determined based on the statistical parameters; and then the score of the first statistical model is determined based on the statistical parameters and the target scene. It can be seen that the data processing method of this application can score the statistical model, and because the scene information of the statistical model is combined in the process of determining the score, the accuracy of the evaluation is also high.

[0082] The following describes the process by which the electronic device determines the target scene corresponding to the first statistical model based at least on the location information in step S202. This process may include, but is not limited to, any one of steps S2021 to S2023 below.

[0083] S2021. If the location information includes a starting point and at least two destinations, the electronic device determines the target scene as a starting point migration scene.

[0084] For example, if the origin includes Suzhou and the destination includes Beijing, Shanghai and Nanjing, then the target scenario is determined as the origin-originating migration scenario.

[0085] S2022. If the location information includes at least two origins and one destination, the electronic device determines the target scene as a destination migration scene.

[0086] For example, if the starting point includes Xi'an, Chengdu, and Taiyuan, and the destination includes Suzhou, then the target scenario is determined as the destination migration scenario.

[0087] S2023. If the location information includes a starting point and a destination; and the statistical parameters also include a first speed and a second speed; then the target scene is determined to be a flow scene between the starting point and the destination.

[0088] For example, if the origin includes Xi'an and the destination includes Suzhou, and the statistical references also include a first speed and a second speed, where the first speed is 150 km / h and the second speed is 400 km / h, then the target scenario is determined to be a destination migration scenario. Specifically, this refers to a scenario where people travel from Xi'an to Suzhou by train or plane.

[0089] The following describes the process by which the S203 electronic device determines the score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scenario.

[0090] In one possible implementation, such as Figure 3 As shown, the process may include, but is not limited to, S2031 to S2034 below.

[0091] S2031. The electronic device determines the completeness of the first statistical model in the target scenario based at least on the statistical time period of the origin and the statistical time period of the destination.

[0092] The completeness of the first statistical model is used to characterize the omissions in the statistical process. Higher completeness means fewer omitted samples and a higher score for the first statistical model.

[0093] The embodiments of this application do not limit the specific implementation process for determining the integrity of the first statistical model, and can be determined according to the actual situation.

[0094] S2032. The electronic device determines the misjudgment rate of the first statistical model in the target scenario, based at least on the statistical time period of the origin and / or the statistical time period of the destination.

[0095] The misclassification rate of the first statistical model is used to characterize the probability of misclassification during the statistical process; that is, the probability of including non-statistical samples in the statistical samples. The higher the misclassification rate, the higher the probability of including non-statistical samples; and the lower the score of the first statistical model.

[0096] The embodiments of this application do not limit the specific implementation process for determining the misclassification rate of the first statistical model, and can be determined according to the actual situation.

[0097] S2033. The electronic device determines the statistical cost of the first statistical model in the target scenario based at least on the statistical time period of the origin and / or the statistical time period of the destination.

[0098] The statistical cost of the first statistical model represents the human and material resources consumed in the statistical process. A higher statistical cost results in a lower score for the first statistical model.

[0099] The embodiments of this application do not limit the specific implementation process for determining the statistical cost of the first statistical model, which can be determined according to the actual situation.

[0100] S2034. The electronic device determines the score of the first statistical model based on the completeness, the false positive rate, and the statistical cost.

[0101] The higher the completeness of the first statistical model, the lower the misjudgment rate of the first statistical model, and the lower the statistical cost of the first statistical model, the higher the score of the first statistical model.

[0102] The embodiments of this application do not impose a unique limitation on the process by which an electronic device determines the score of the first statistical model based on the completeness of the first statistical model, the misjudgment rate of the first statistical model, and the statistical cost of the first statistical model; the process can be determined according to the actual situation.

[0103] Understandably, more factors can be considered when determining the score of the first statistical model; alternatively, only some factors such as completeness, false positive rate, and statistical cost can be considered. For specific implementation details, please refer to the detailed descriptions in S2031 to S2034 above, which will not be repeated here.

[0104] The following describes the process by which the S2034 electronic device determines the score of the first statistical model based on the integrity, the false positive rate, and the statistical cost.

[0105] In one possible implementation, the process may include: determining the score of the first statistical model using a first formula.

[0106] The first formula includes: Wherein, effect represents the score of the first statistical model; α represents the completeness of the first statistical model; β represents the false positive rate of the first statistical model; and γ represents the statistical cost of the first statistical model.

[0107] In another possible implementation, when determining the score of the first statistical model, different weights can be assigned to completeness, misclassification rate and statistical cost, that is, different coefficients can be assigned to α, β and γ in the first formula. The specific implementation process will not be described here.

[0108] The following describes the process by which the S2031 electronic device determines the completeness of the first statistical model in the target scenario based at least on the statistical time period of the origin and the statistical time period of the destination. This process may include, but is not limited to, Case 1 or Case 2 below.

[0109] Case 1: When the target scenario is a migration from the origin or a migration to the destination, determine the completeness of the first statistical model;

[0110] Case 2: When the target scenario is a flow between the origin and destination, determine the completeness of the first statistical model.

[0111] The following section explains the process of determining the completeness of the first statistical model when the target scenario is either a migration out of the origin or a migration into the destination.

[0112] Case 1 may include: if the target scenario is a migration out of origin scenario or a migration in destination scenario, the electronic device determines that the completeness α of the first statistical model is N2-N1 under the target scenario.

[0113] Wherein, α represents the completeness of the first statistical model, N2 represents the duration between the statistical end time of the destination and the statistical end time of the origin, and N1 represents the duration between the statistical start time of the destination and the statistical start time of the origin.

[0114] Understandable, it can be done through T D -T O This indicates the completeness of the first statistical model. Where T... DT represents the statistical duration of the destination. O Indicates the statistical duration at the starting point.

[0115] The following section explains the process of determining the completeness of the first statistical model when the target scenario in Case 2 is a flow scenario between the origin and the destination.

[0116] In scenario 2, the statistical parameters of the first statistical model also include a first velocity, a second velocity, and a first distance. The first velocity is less than the second velocity.

[0117] Case 2 may include: the electronic device determines the completeness of the first statistical model using a second formula.

[0118] The second formula includes: α represents the completeness of the first statistical model, and T O The T represents the statistical duration of the origin. D The statistical time to the destination is represented by L, which represents the first distance between the origin and the destination, V1 represents the first speed, and V2 represents the second speed.

[0119] The following describes the process by which the S2032 electronic device determines the misclassification rate of the first statistical model in the target scenario, based at least on the statistical time period of the origin and / or the statistical time period of the destination. This process may include, but is not limited to, the following situation A or situation B.

[0120] Case A: When the target scenario is a migration from the origin or a migration to the destination, determine the misclassification rate of the first statistical model;

[0121] In Scenario B, when the target scenario is a flow between the origin and destination, determine the misclassification rate of the first statistical model.

[0122] The following explains the process of determining the misclassification rate of the first statistical model when the target scenario is either a migration out of the origin or a migration into the destination.

[0123] Scenario A may include: the electronic device determines the misclassification rate of the first statistical model using a third formula;

[0124] The third formula includes: Wherein, β represents the misclassification rate of the first statistical model, and S O The S represents the statistical standard for the length of stay at the origin. D The T represents the statistical standard for the length of stay at the destination. O The T represents the statistical duration of the origin. D This indicates the statistical duration of the destination.

[0125] The following section explains the process of determining the misclassification rate of the first statistical model in case B, where the target scenario is a flow between the origin and destination.

[0126] In scenario B, the statistical parameters of the first statistical model also include a first velocity, a second velocity, and a first distance. The first velocity is less than the second velocity.

[0127] Scenario B may include: the electronic device determines the misclassification rate of the first statistical model using the fourth formula.

[0128] The fourth formula includes: β represents the misclassification rate of the first statistical model, and T D The statistical time to the destination is represented by L, which represents the first distance between the origin and the destination, V1 represents the first speed, and V2 represents the second speed.

[0129] The following describes the process by which the electronic device determines the statistical cost of the first statistical model in the target scenario, based at least on the statistical time period of the origin and / or the statistical time period of the destination, in step S2033. This process may include, but is not limited to, any one of the following methods 1 to 3.

[0130] Method 1: When the target scenario is a migration from the origin, determine the statistical cost of the first statistical model;

[0131] Method 2: When the target scenario is a destination migration scenario, determine the statistical cost of the first statistical model;

[0132] Method 3: When the target scenario is a flow scenario between the origin and the destination, determine the statistical cost of the first statistical model.

[0133] For method 1, if the target scenario is a migration from the origin, the electronic device determines that the statistical cost of the first statistical model in the target scenario is T. D .

[0134] The T D This indicates the statistical duration of the destination.

[0135] Understandably, the statistical cost of the first statistical model can be determined to be T. D ×n1; where n1 represents the number of destinations.

[0136] For method 2, if the target scenario is a destination migration scenario, the electronic device determines that the statistical cost of the first statistical model in the target scenario is T. O .

[0137] The T OIndicates the statistical duration at the starting point.

[0138] It is understandable, and the statistical cost of the first statistical model can be determined to be T. D ×n2; where n2 represents the number of starting locations.

[0139] For method 3, if the target scenario is a flow scenario between the origin and the destination, the electronic device determines that the statistical cost of the first statistical model in the target scenario is T. D +T O .

[0140] It is understandable, and the statistical cost of the first statistical model can be determined to be T. D .

[0141] The data processing method provided in this application embodiment can also adjust the first statistical model based on the score of the first statistical model. For example... Figure 4 As shown, the specific adjustment process may include, but is not limited to, S204 and S205 below.

[0142] S204. The electronic device determines that the score of the first statistical model is less than the first score threshold.

[0143] The first scoring threshold is used to determine whether the first statistical model needs adjustment. Specifically, if the score of the first statistical model is less than the first scoring threshold, it means that the first statistical model needs adjustment; if the score of the first statistical model is greater than or equal to the first scoring threshold, it means that the first statistical model does not need adjustment.

[0144] The embodiments of this application do not limit the value of the first scoring threshold, which can be determined according to the actual situation.

[0145] S205. The electronic device adjusts the adjustment parameters of the first statistical model.

[0146] The adjustment parameters include one or more of the following: the statistical period of the origin, the statistical period of the destination, the statistical standard for the length of stay at the origin, and the statistical standard for the length of stay at the destination.

[0147] S205 can be implemented as follows: the electronic device adjusts one or more of the adjustment parameters of the first statistical model to increase the score of the adjusted first model; after multiple adjustments, a first statistical model with a score greater than the first score threshold is obtained.

[0148] The following example, using population mobility as an example, illustrates the data processing method provided in this application.

[0149] With the development of urban transportation and the market economy, population mobility has become very frequent, and research on population mobility phenomena has gradually increased, including exploring the purpose of population mobility, conducting population migration statistics in special scenarios in a certain way, and predicting population mobility trends.

[0150] For example, related technologies propose a method for determining the purpose of population migration, which mainly includes: if a user is detected to be migrating, then determining the user's stop point based on the user's trajectory point data in the destination region; determining the user's stop point of interest based on the user's stop point; and then determining the user's purpose of migration based on the user's stop point of interest.

[0151] For example, a related technology proposes a method for identifying population migration based on mobile phone signaling. This method primarily involves: using mobile phone signaling data as the basis for identifying user migration has the advantages of a large user base, wide coverage, and dynamic, continuous data. Therefore, by continuously tracking and monitoring users' residential locations over multiple months, the monthly residential locations are converted into a set of spatial locations with temporal characteristics. Then, by establishing a temporal spatial data clustering model, the method automatically identifies people whose residences are migrating, records the time of migration, and information such as the entry / exit locations, thereby dynamically understanding the patterns of population migration.

[0152] For example, related technologies propose a method for predicting the scale of returning populations from epidemic areas based on big data on population migration. This method mainly includes: obtaining the inflow scale index and the proportion of origin of those migrating into epidemic area j during the Spring Festival travel rush from Baidu Maps migration big data, as well as obtaining the resident population of epidemic area j; obtaining the scale of the population migrating from target region i into epidemic area j during the Spring Festival travel rush based on the inflow scale index, the proportion of origin, and the resident population of epidemic area j; obtaining the scale of the non-returning population migrating from target region i into epidemic area j during the Spring Festival travel rush; and obtaining the scale of the returning population after the Spring Festival based on the population scale and the non-returning population scale.

[0153] For example, related technologies propose a method for estimating the number of infected people in an epidemic based on big data on population migration. This method mainly includes: obtaining the number of people migrating from each epidemic city to each target city, the infection rate in each epidemic city, and the resident population, resident population, main road length, railway length, and number of residential areas in each target city / county; using the infection rate in each epidemic city as a weighting coefficient, calculating a weighted sum of the population to obtain a first weighted population; obtaining a second weighted population based on the first weighted population; and inputting the second weighted population, the resident population, main road length, railway length, and number of residential areas in each target district / county into a trained epidemic infection estimation model to obtain the number of infected people in each target district / county.

[0154] Analysis of the relevant technologies reveals the following main drawbacks:

[0155] 1. No population statistical model has been formed.

[0156] The methods for determining the purpose of population migration only target the identification of individual users' migration behavior and have not formed a clear group statistical model, making it difficult to directly form effective large-scale and social statistical conclusions from the identification of each individual's migration behavior.

[0157] 2. Limited to a single statistical method

[0158] The method for identifying population migration based on mobile signaling, which studies population mobility, is limited to using a certain statistical method (mobile base station signaling statistics) to obtain changes in users' residences, and does not specify a group statistical model.

[0159] 3. Limited to a specific application scenario.

[0160] The methods for predicting the scale of people returning from epidemic areas based on big data of population migration and the methods for estimating the number of people infected in epidemics based on big data of population migration are both designed for the scenario of epidemic prevention and control, and are used to predict the population migration statistics and flow trends in special scenarios. Since the models incorporate the characteristics of epidemic transmission, the above-mentioned population flow scenarios may not be applicable to scenarios such as daily office commuting and migrant workers returning home.

[0161] It can be seen that there are many studies on population flow models for specific scenarios, but it is difficult to compare and evaluate the value and expected differences of the statistical results of various models in different dimensions.

[0162] To address the aforementioned shortcomings, this embodiment of the application solves the following problems: establishing an evaluation system for population mobility models, assessing the functional effectiveness of population mobility statistical models based on this evaluation system, and helping researchers select appropriate models and methods for population mobility statistics and prediction when solving population mobility problems in different scenarios. In population mobility analysis research addressing multiple scenario needs, the parameters of the population mobility model can be adjusted according to this evaluation system to obtain controllable, matching, and convincing statistical results.

[0163] The specific details of this embodiment will now be described.

[0164] The main principles include: This embodiment of the application provides an evaluation system and optimization method for a multi-scenario population flow model (equivalent to a first statistical model). By analyzing the structure and factors of various population flow models, common population flow model parameters (equivalent to statistical parameters) are identified. For the common population flow model parameters, a common evaluation system for the population flow model is formed, and based on the evaluation system, the effectiveness of a certain population flow model in a certain application scenario is evaluated. For different population flow statistical application scenarios, the model parameters are adjusted to obtain the optimal population flow statistical model.

[0165] The key parameters and scenarios of the general population mobility model are explained below.

[0166] Key parameters include: time, location, and people.

[0167] Location parameters can include: origin (O) and destination (D).

[0168] The distance between the center point of the starting point and the center point of the destination is L.

[0169] Regarding time parameters:

[0170] From an individual perspective, the duration of stay at the place of origin is recorded as O. t The statistical period for the destination is denoted as D. t .

[0171] Among them, O t This can be expressed by the following formula 1; D t It can be expressed by the following formula 2.

[0172]

[0173] In Formula 1, O t This indicates the duration of an individual's stay at the place of origin. Indicates the start time of an individual's stay at the origin. This indicates the end time of an individual's stay at the origin.

[0174]

[0175] In Formula 1, D t This indicates the duration of an individual's stay at the destination. Indicates the start time of an individual's stay at the destination; This indicates the end time of an individual's stay at the destination.

[0176] Therefore, the O of a person (individual) in location O. t The duration of stay within a given time period is denoted as t.O To place someone (individual) in location D t The duration of stay within a given time period is denoted as t. D .

[0177] For the dwell time of each individual, a statistical standard value is set as S. O and S D S O S represents the standard for individual length of stay in location O. D This represents the standard for the duration of an individual's stay in location D. That is, if an individual's stay in location O is less than S... O If an individual's stay in location D is less than that of location S, then that individual will not be included in the statistics; D If the individual is not included in the statistics, then that individual will not be included.

[0178] From the perspective of statistical population, the statistical period of the place of origin is denoted as O. T The statistical duration is T. O The statistical period at the destination is denoted as D. T The statistical duration is T. D Among them, O T T can be expressed by the following formula 3, D It can be represented by the following formula 4.

[0179]

[0180] In formula 3, O T Indicates the statistical period of the group at its origin; Indicates the start time of the statistics for the group at the origin; This indicates the end time of the statistics for the group at the starting point.

[0181]

[0182] In formula 4, D T Indicates the statistical period of the group at the destination; Indicates the start time of the group's statistics at the destination. This indicates the end time of the group's statistics at the destination.

[0183] The time parameters satisfy the following relationship:

[0184]

[0185]

[0186]

[0187]

[0188] Where N1 represents the duration between the statistical start time of the group at the destination and the statistical start time of the group at the origin; N2 represents the duration between the statistical end time of the group at the destination and the statistical end time of the group at the origin.

[0189] S O ∈(0,T) O ];S D ∈(0,T) D ];

[0190] People: The target population for statistical analysis (P), denoted as p, is a single individual. Each individual is included in the statistics only if: t O ≥S O And t D ≥S D .

[0191] The following is a description of the scenario.

[0192] Specifically, this may include, but is not limited to, the following scenarios 1 to 3.

[0193] Scenario 1: Taking location O as the research object, study the problem of tracking the migration of people from location O.

[0194] For example, analyzing the employment destinations of local university graduates and issues related to the control of the spread of local infectious diseases.

[0195] Scenario 2: Taking location D as the research object, study the source of population migration to location D.

[0196] For example, investigating the origins of tourists to local attractions and addressing issues related to the management and control of visitors from outside the area.

[0197] Scenario 3: Taking the flow path (L) between ODs as the research object, study the flow mode.

[0198] For example, issues related to predicting flight and train schedules during the Spring Festival travel rush.

[0199] The population mobility statistical model will be explained below.

[0200] Different population flow statistics models can be obtained depending on the different scenarios, including but not limited to the following Scenario 1 to Scenario 3 models.

[0201] Population flow model for scenario 1:

[0202] O T The time period is in location O, and the length of stay meets the S requirement. O ;D T The time period is not in location O, and the length of stay in all potential locations D satisfies S. DThe number of people, and a statistical ranking of all potential D locations based on the above number of people.

[0203] Population flow model for scenario 2:

[0204] Statistics D T The time period is in location D, and the length of stay satisfies S. D ;O T The time period is not in location D, and the length of stay in all potential locations O satisfies S. O The number of people, and a statistical ranking of all potential D locations based on the above number of people.

[0205] Population flow model for scenario 3:

[0206] Statistics O T The time period is in location O, and D T The population flow statistics are based on the number of people in location D during a given time period, provided that the speed of movement for a certain mode of transportation is V. The movement behavior of different individuals can be grouped and statistically analyzed according to their movement speed between location O and location D, thus yielding population statistics for different modes of transportation.

[0207] The following section explains the evaluation system for multi-scenario population mobility models.

[0208] This embodiment of the application targets a population mobility model and has three evaluation factors: statistical completeness α, statistical misjudgment rate β, and statistical cost γ.

[0209] Among them, the larger the statistical completeness α value, the more comprehensive the target population that should be included in the statistics; the larger the statistical misjudgment rate β, the greater the probability of non-target population in the statistical results; and the smaller the statistical cost γ, the smaller the statistical expense.

[0210] Therefore, the benefit of a population mobility model can be expressed as: statistical completeness per unit cost minus the misjudgment rate. For example, it can be expressed as Equation 5 below (equivalent to Equation 1 above).

[0211]

[0212] In Formula 5, effect represents the benefit of a population mobility model; α represents the statistical completeness of the population mobility model; β represents the statistical misjudgment rate of the population mobility model; and γ represents the statistical cost of the population mobility model.

[0213] Among them, the population inflow process model for a specific target area can be as follows: Figure 5 As shown, from Figure 5 It can be seen that S O T O S D T D The relationship between N1 and N2.

[0214] The following section explains the process of determining the benefits of population mobility models for specific scenarios.

[0215] For scenario 1, “Taking location O as the research object, studying the problem of tracking the migration of people from location O”, the process of determining the effectiveness of the population flow model may include, but is not limited to, the following S11 to S14.

[0216] S11. Determine the completeness of the population flow model for scenario 1.

[0217] Considering statistical completeness α: Comparing the effects of increasing or decreasing values ​​of the above parameters, we find that the smaller N1 is, the more timely and thorough the statistical requirements for emigrant populations are, meaning that the problems are addressed as soon as they arise. The outflow population statistics should be initiated immediately after the event; the larger N2 is, the more thorough and accurate the outflow population statistics are, meaning that the statistics are accurate even after the event has ended. For a long time afterward, the impact of this event on population movement was continuously observed.

[0218] Therefore, N2-N1 or T can be used. D -T O This represents statistical completeness.

[0219] S12. Determine the misjudgment rate of the population flow model in scenario 1.

[0220] Considering the statistical false positive rate β: from the perspective of not missing any suspicious persons, S O / T O and S D / T D The smaller the value of S, the more comprehensive the population count. When S approaches zero, people from all sources will be counted, increasing the probability that unrelated individuals passing by will be mistakenly included in the count. Conversely, when S... O / T O and S D / T D When all values ​​tend to be 1, the residency conditions included in the statistics become more stringent, but the false positive rate will be greatly reduced.

[0221] Therefore, 1-(S) can be used. O ×S D ) / (T O ×T D (Equivalent to the third formula above) represents the statistical misjudgment rate.

[0222] S13. Determine the cost of the population flow model for scenario 1.

[0223] Considering the statistical cost γ: as T DAs the number of events increases, the monitoring time for the target statistical events becomes longer, the amount of statistical data increases, and the statistical cost also increases.

[0224] Therefore, T can be used. D This represents the cost of statistics.

[0225] S14. Determine the benefits of the population mobility model in scenario 1.

[0226] For scenario 1, “Taking location O as the research object, studying the problem of tracking the migration of people from location O”, the formula for calculating the population mobility benefits can be shown in Formula 6 below.

[0227]

[0228] In Formula 6, Effect1 represents the benefit of the population flow model corresponding to Scenario 1, S O S represents the standard for individual length of stay in location O. D T represents the standard duration of individual stay in location D. O T represents the statistical duration of a group in location O. D This indicates the statistical duration of the group in location D.

[0229] Similar to scenario 1, for scenario 2, "taking location D as the research object and studying the source of population migration to location D", the formula for calculating population mobility benefits can be shown in Formula 7 below.

[0230]

[0231] In Formula 7, Effect2 represents the benefit of the population flow model corresponding to Scenario 2, S O S represents the standard for individual length of stay in location O. D T represents the standard duration of individual stay in location D. O T represents the statistical duration of a group in location O. D This indicates the statistical duration of the group in location D.

[0232] For scenario 3, “taking the flow path between OD as the research object and studying the mode of population flow”, the process of determining the benefits of the population flow model may include, but is not limited to, the following S21 to S24.

[0233] S21. Determine the completeness of the population flow model for scenario 2.

[0234] Considering statistical completeness α: Given the distance L between locations O and D, if we want to count the number of people whose population movement speed between O and D falls within the range [V1, V2], that is, the movement time of each individual should be between [L / V2, L / V1]. Based on this, we can derive the following: T For the target time period, when T D =T OWhen +L / V1-L / V2, D T The effective statistical period can fully cover O T The outflowing population is recorded as having a statistical completeness of 1.

[0235] Then when D T The effective statistical period is any value T. D When the statistical completeness α is T D / (T O +L / V1-L / V2).

[0236] Figure 6 It can be seen that O T D T L / V min L / V max The relationship between them; Figure 6 In the diagram, long dashed lines represent misjudgments; short dashed lines represent non-misjudgments.

[0237] S22. Determine the misjudgment rate of the population flow model in scenario 2.

[0238] Considering the statistical misjudgment rate β: From a statistical perspective, this scenario does not concern itself with the duration (t) of each statistical individual's stay in locations O and D. O and t D (Anything greater than 0 is acceptable), meaning there's no need to consider the influence of S. O and S D The error rate can be high when the value is very small.

[0239] However, when counting the number of people whose flow rate is between V1 and V2, the error was mistakenly... When statistically analyzing a group of motions at a certain velocity between [V2, ∞] and [V2, ∞], it's possible to mistakenly include some groups moving at even faster speeds. It can be observed that as V2 increases, the two misclassified intervals shrink.

[0240] when At that time, T D =L / V1-L / V2 has the smallest misjudged speed range, T D The range of speeds that will be misjudged will increase as the speed is increased.

[0241] Remember T D When β = L / V1 - L / V2, the false positive rate is 0, so β ​​can be expressed as 1 - (L / V1 - L / V2) / T D .

[0242] S23. Determine the cost of the population flow model for scenario 2.

[0243] Considering the statistical cost γ: as T D As the number of events increases, the monitoring time for the target statistical events becomes longer, the amount of statistical data increases, and the statistical cost also increases.

[0244] Therefore, using T D This represents the cost of statistics.

[0245] S24. Determine the benefits of the population flow model in scenario 2.

[0246] For scenario 3, "taking the flow path between OD as the research object and studying the mode of population flow", the population flow benefit calculation formula can be shown in the following formula 8.

[0247]

[0248] In Formula 8, Effect3 represents the benefit of the population flow model corresponding to Scenario 3, and T O T represents the statistical duration of a group in location O. D V1 represents the statistical duration of the population at location D; L represents the distance between O and D; V1 represents the minimum statistical flow velocity; and V2 represents the maximum statistical flow velocity.

[0249] It can be seen that by effectively controlling the above parameter values ​​in the population flow behavior model between two places, the statistical integrity, statistical cost and misjudgment rate of the target model can be evaluated and optimized.

[0250] The evaluation and optimization process of the population mobility model will be explained below.

[0251] like Figure 7 As shown, the specific implementation may include, but is not limited to, S31 to S33 below.

[0252] S31. Determine the target area and select the scene.

[0253] The selection of scenarios can be based on the general analytical structure of population flow models.

[0254] If the target area is O, it is Scene 1; if the target area is D, it is Scene 2; if the target area is both O and D, it is Scene 3.

[0255] For example, if we take location O as the research object and study the problem of tracking the migration out of location O, we choose scenario 1; if we take location D as the research object and study the problem of the source of migration into location D, we choose scenario 2; if we take the flow path (L) between O and D as the research object and study the flow mode, we choose scenario 3.

[0256] S32. Clarify the analysis constraints in the current scenario.

[0257] In both Scenario 1 and Scenario 2, the duration of statistical events for the population outflow location (location O) and the population inflow location (location D) needs to be clearly defined. That is, the time period during which population outflow occurs in the statistical events for the outflow location and the time period during which population inflow occurs in the statistical events for the inflow location need to be clearly defined.

[0258] In scenario 3, it is necessary to define the speed range of population movement to be counted in order to count the number of people with homogeneous movement behavior characteristics within that speed range.

[0259] The analytical constraints determined in this step are readily available in each scenario and are controllable based on practical significance. Subsequent statistical results should be used to build models for more subdivided scenarios under the above constraints.

[0260] S33. Adjust the model parameters using the scenario benefit formula and select the model parameters with the highest benefit.

[0261] Based on the above scenarios, different formulas for calculating population mobility benefits can be selected, and the results can be used to optimize variable parameters. Choosing an evaluation model with a high benefit value is beneficial for improving the statistical effect per unit statistical cost. Furthermore, based on the aforementioned benefit calculation formulas, the parameter values ​​in the model can be effectively adjusted to optimize the model in the direction of increasing benefits.

[0262] For scenario 1, the dwell time at location O, the dwell time at location D, and the statistical time at location D can be adjusted;

[0263] For scenario 2, the dwell time at location O, the dwell time at location D, and the statistical time at location O can be adjusted;

[0264] For scenario 3, the statistical time periods for location O and location D can be adjusted;

[0265] This embodiment of the application has the following characteristics:

[0266] First, this embodiment provides a general analytical structure for population flow models.

[0267] Including key parameters such as time, location, and people, three types of research scenarios and corresponding population flow models are: taking location O as the research object, studying the problem of tracking the migration out of location O; taking location D as the research object, studying the problem of the source of migration into location D; taking the flow path (L) between O and D as the research object, studying the flow behavior patterns.

[0268] Secondly, this embodiment provides a multi-scenario population flow model evaluation system.

[0269] There are three evaluation factors: statistical integrity α, statistical cost β, and statistical misjudgment rate γ. Combining these three evaluation factors, a benefit formula for a general population mobility model suitable for multiple scenarios is derived, which is the statistical integrity minus the misjudgment rate under unit cost.

[0270] Thirdly, this embodiment provides a method for optimizing the implementation of a multi-scenario population flow model.

[0271] This paper addresses three scenarios of population mobility: "study of population emigration from location O," "study of population inflow from location D," and "population mobility between two locations." It discusses statistical completeness and potential misjudgments in each scenario, and establishes specific methods for calculating the benefits of different population mobility models based on known constraints. By adjusting the statistical parameters for each scenario, a population mobility model suitable for the specific constraints and objectives can be obtained.

[0272] This embodiment of the application has the following technical effects:

[0273] First, a universal statistical evaluation system for population mobility has been established.

[0274] This embodiment is based on a general population flow model, listing all scenarios from both ends to the middle of the path for the research subjects. The model is evaluated from a statistical perspective. The evaluation effect is not affected by individual characteristics. It takes into account statistical completeness and statistical error rate, and also incorporates statistical costs, so that the evaluation results are more in line with the actual economic market characteristics.

[0275] Second, it is not limited to a specific application scenario or field.

[0276] This embodiment evaluates parameters based on three types of characteristic scenario models in the population mobility model, and finds evaluation influencing factors of different dimensions for different characteristic scenarios.

[0277] It can be seen that the same set of population flow data will yield different statistical results depending on the research scenario: "analyzing the characteristics of outflowing population based on the origin," "analyzing the characteristics of inflowing population based on the destination," or "analyzing the flow behavior between two locations based on the flow route." Furthermore, the parameters of the population flow model can be adjusted according to the requirements of the evaluation indicators in different scenarios to obtain controllable, suitable, and convincing statistical results.

[0278] Third, it is not limited to a single statistical method.

[0279] The population flow model in this embodiment does not limit the data collection methods; the data collection methods can be any means. For example, mobile phone signaling information, GPS location information, or cameras installed at the origin and destination of the flow, etc.

[0280] Secondly, to implement the above-mentioned data processing method, an embodiment of this application provides a data processing apparatus, which will be described below in conjunction with... Figure 8 The structural diagram of the data processing device shown is used for illustration.

[0281] like Figure 8 As shown, the data processing device 80 includes: an acquisition unit 801, a first determination unit 802, and a second determination unit 803. Wherein:

[0282] The obtaining unit 801 is used to obtain statistical parameters of the first statistical model; the statistical parameters include at least: location information, statistical time period of the origin, and statistical time period of the destination;

[0283] The first determining unit 802 is used to determine the target scene corresponding to the first statistical model based at least on the location information;

[0284] The second determining unit 803 is used to determine the score of the first statistical model based at least on the statistical time period of the origin, the statistical time period of the destination, and the target scenario; the score of the first statistical model is used to evaluate the first statistical model.

[0285] In some embodiments, the first determining unit 802 is specifically used for:

[0286] If the location information includes one origin and at least two destinations, then the target scenario is determined to be an origin-origin migration scenario.

[0287] If the location information includes at least two origins and one destination, then the target scenario is determined to be a destination migration scenario.

[0288] If the location information includes a starting point and a destination, and the statistical parameters also include a first speed and a second speed, then the target scenario is determined to be a flow scenario between the starting point and the destination.

[0289] In some embodiments, the second determining unit 803 is specifically used for:

[0290] The completeness of the first statistical model in the target scenario is determined based at least on the statistical time period of the origin and the statistical time period of the destination.

[0291] Based at least on the statistical time period of the origin and / or the statistical time period of the destination, determine the misjudgment rate of the first statistical model in the target scenario;

[0292] The statistical cost of the first statistical model in the target scenario is determined based at least on the statistical time period of the origin and / or the statistical time period of the destination.

[0293] The score of the first statistical model is determined based on the completeness, the false positive rate, and the statistical cost.

[0294] In some embodiments, the second determining unit 803 is further configured to:

[0295] The score of the first statistical model is determined by the first formula;

[0296] The first formula includes: The effect represents the score of the first statistical model; α represents the completeness of the first statistical model; β represents the false positive rate of the first statistical model; and γ represents the statistical cost of the first statistical model.

[0297] In some embodiments, the second determining unit 803 is further configured to:

[0298] If the target scenario is a migration out of origin or a migration into destination, the completeness of the first statistical model is determined to be α = N2 - N1 under the target scenario; where α represents the completeness of the first statistical model, N2 represents the duration between the statistical end time of the destination and the statistical end time of the origin, and N1 represents the duration between the statistical start time of the destination and the statistical start time of the origin.

[0299] If the target scenario is a flow scenario between a starting point and a destination, the statistical parameters further include a first speed, a second speed, and a first distance, wherein the first speed is less than the second speed; the completeness of the first statistical model is determined by a second formula; wherein the second formula includes: α represents the completeness of the first statistical model, and T O The T represents the statistical duration of the origin. D The statistical time to the destination is represented by L, which represents the first distance between the origin and the destination, V1 represents the first speed, and V2 represents the second speed.

[0300] In some embodiments, the second determining unit 803 is further configured to:

[0301] If the target scenario is a migration out of origin or a migration in destination, the misclassification rate of the first statistical model is determined by a third formula; wherein, the third formula includes: Wherein, β represents the misclassification rate of the first statistical model, and SO The S represents the statistical standard for the length of stay at the origin. D The T represents the statistical standard for the length of stay at the destination. O The T represents the statistical duration of the origin. D Indicates the statistical duration of the destination;

[0302] If the target scenario is a flow scenario between a starting point and a destination; the statistical parameters also include a first speed, a second speed, and a first distance, wherein the first speed is less than the second speed; the misclassification rate of the first statistical model is determined by a fourth formula; wherein the fourth formula includes: β represents the misclassification rate of the first statistical model, and T D The statistical time to the destination is represented by L, which represents the first distance between the origin and the destination, V1 represents the first speed, and V2 represents the second speed.

[0303] In some embodiments, the second determining unit 803 is further configured to:

[0304] If the target scenario is a migration from the origin, then the statistical cost of the first statistical model in the target scenario is determined to be T. D The T D Indicates the statistical duration of the destination;

[0305] If the target scenario is a destination migration scenario, then the statistical cost of the first statistical model under the target scenario is determined to be T. O The T O Indicates the statistical duration at the origin;

[0306] If the target scenario is a flow between a starting point and a destination, then the statistical cost of the first statistical model in the target scenario is determined to be T. D +T O .

[0307] In some embodiments, the data processing apparatus 80 may further include an adjustment unit, which is configured to perform:

[0308] It is determined that the score of the first statistical model is less than the first score threshold;

[0309] Adjust the adjustment parameters of the first statistical model; the adjustment parameters include one or more of the following: the statistical period of the origin, the statistical period of the destination, the statistical standard of the length of stay at the origin, and the statistical standard of the length of stay at the destination.

[0310] It should be noted that the data processing device provided in this application embodiment includes all the units included, which can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field-programmable gate array (FPGA), etc.

[0311] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0312] It should be noted that, in the embodiments of this application, if the above-described data processing method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0313] Thirdly, to implement the above data processing method, embodiments of this application provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements the steps in the data processing method provided in the above embodiments.

[0314] The following is combined Figure 9 The electronic device 90 shown is illustrated with a structural diagram of the electronic device.

[0315] In one example, electronic device 90 can be the aforementioned electronic device. For example... Figure 9As shown, the electronic device 90 includes: a processor 901, at least one communication bus 902, a user interface 903, at least one external communication interface 904, and a memory 905. The communication bus 902 is configured to enable communication between these components. The user interface 903 may include a display screen, and the external communication interface 904 may include standard wired and wireless interfaces.

[0316] The memory 905 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or already processed by the processor 901 and various modules in the electronic device (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).

[0317] Fourthly, embodiments of this application provide a storage medium, namely a computer-readable storage medium, on which a computer program is stored, which, when executed by a processor, implements the steps in the data processing method provided in the above embodiments.

[0318] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0319] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0320] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0321] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0322] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0323] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0324] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0325] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0326] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the statistical parameters of the first statistical model; the statistical parameters include at least: location information, the statistical time period of the origin, and the statistical time period of the destination; Based at least on the location information, the target scene corresponding to the first statistical model is determined; Based at least on the statistical time period of the origin and the statistical time period of the destination, determine the completeness of the first statistical model in the target scenario; based at least on the statistical time period of the origin and / or the statistical time period of the destination, determine the misclassification rate of the first statistical model in the target scenario; based at least on the statistical time period of the origin and / or the statistical time period of the destination, determine the statistical cost of the first statistical model in the target scenario; determine the score of the first statistical model through a first formula; the first formula includes: The This represents the score of the first statistical model; the This indicates the completeness of the first statistical model; the This represents the misclassification rate of the first statistical model; the The first statistical model represents the statistical cost; the score of the first statistical model is used to evaluate the first statistical model.

2. The method according to claim 1, characterized in that, Determining the target scene corresponding to the first statistical model based at least on the location information includes: If the location information includes one origin and at least two destinations, then the target scenario is determined to be an origin-origin migration scenario. If the location information includes at least two origins and one destination, then the target scenario is determined to be a destination migration scenario. If the location information includes a starting point and a destination, and the statistical parameters also include a first speed and a second speed, then the target scenario is determined to be a flow scenario between the starting point and the destination.

3. The method according to claim 1, characterized in that, Determining the completeness of the first statistical model in the target scenario, based at least on the statistical time period of the origin and the statistical time period of the destination, includes: If the target scenario is a migration from origin to destination or a migration to destination, determine the completeness of the first statistical model under the target scenario. ; wherein, the This indicates the completeness of the first statistical model, the This indicates the duration between the statistical end time of the destination and the statistical end time of the origin. This indicates the duration between the statistical start time of the destination and the statistical start time of the origin. If the target scenario is a flow scenario between a starting point and a destination, the statistical parameters further include a first speed, a second speed, and a first distance, wherein the first speed is less than the second speed; the completeness of the first statistical model is determined by a second formula; wherein the second formula includes: The This indicates the completeness of the first statistical model, the The statistical duration indicating the origin, the The statistical duration of the destination, the The first distance between the origin and the destination is represented by the following: Indicates the first speed, the This indicates the second speed.

4. The method according to claim 1, characterized in that, Determining the misclassification rate of the first statistical model in the target scenario, based at least on the statistical time period of the origin and / or the statistical time period of the destination, includes: If the target scenario is a migration out of origin or a migration in destination, the misclassification rate of the first statistical model is determined by a third formula; wherein, the third formula includes: ; wherein, the This represents the misclassification rate of the first statistical model, the The statistical standard for the length of stay at the origin is indicated. The statistical standard for the length of stay at the destination is described in the following text. The statistical duration indicating the origin, the Indicates the statistical duration of the destination; If the target scenario is a flow scenario between a starting point and a destination; the statistical parameters also include a first speed, a second speed, and a first distance, wherein the first speed is less than the second speed; the misclassification rate of the first statistical model is determined by a fourth formula; wherein the fourth formula includes: The This represents the misclassification rate of the first statistical model, the The statistical duration of the destination, the The first distance between the origin and the destination is represented by the following: Indicates the first speed, the This indicates the second speed.

5. The method according to claim 1, characterized in that, Determining the statistical cost of the first statistical model in the target scenario, based at least on the statistical time period of the origin and / or the statistical time period of the destination, includes: If the target scenario is a migration from the origin, then the statistical cost of the first statistical model under the target scenario is determined to be... The Indicates the statistical duration of the destination; If the target scenario is a destination migration scenario, then the statistical cost of the first statistical model under the target scenario is determined to be: The Indicates the statistical duration at the origin; If the target scenario is a flow scenario between the origin and the destination, then the statistical cost of the first statistical model under the target scenario is determined to be: .

6. The method according to claim 1, characterized in that, The method further includes: It is determined that the score of the first statistical model is less than the first score threshold; Adjust the adjustment parameters of the first statistical model; the adjustment parameters include one or more of the following: the statistical period of the origin, the statistical period of the destination, the statistical standard of the length of stay at the origin, and the statistical standard of the length of stay at the destination.

7. A data processing apparatus, characterized in that, The device includes: The obtaining unit is used to obtain statistical parameters of the first statistical model; the statistical parameters include at least: location information, statistical time period of the origin, and statistical time period of the destination; The first determining unit is configured to determine the target scene corresponding to the first statistical model based at least on the location information. The second determining unit is configured to: determine the completeness of the first statistical model in the target scenario based at least on the statistical time period of the origin and the statistical time period of the destination; determine the misclassification rate of the first statistical model in the target scenario based at least on the statistical time period of the origin and / or the statistical time period of the destination; determine the statistical cost of the first statistical model in the target scenario based at least on the statistical time period of the origin and / or the statistical time period of the destination; and determine the score of the first statistical model using a first formula; the first formula includes: The This represents the score of the first statistical model; the This indicates the completeness of the first statistical model; the This represents the misclassification rate of the first statistical model; the The first statistical model represents the statistical cost; the score of the first statistical model is used to evaluate the first statistical model.

8. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the program to implement the data processing method of any one of claims 1 to 6.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data processing method according to any one of claims 1 to 6.