Target object description fusion framework and device under size data, equipment and medium

By combining the fusion framework of small data and big data in social simulation, the positioning and representativeness of social media user portraits in social simulation is solved, a more comprehensive population portrait is achieved, and more effective social simulation is supported.

CN120162480APending Publication Date: 2025-06-17北京大学武汉人工智能研究院
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510133470.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to use directly in social simulations through social media user portraits, especially with challenges in positioning and representation.

Method used

A target object description fusion framework under large and small data is proposed. The target community is selected through electronic maps, combined with active surveys and mobile Internet data, and distributed fusion algorithms and weight parameter optimization are adopted to achieve the fusion of small data and big data.

Benefits of technology

It has achieved a more comprehensive portrayal of people, effectively supported social simulation, and overcome the problems of imbalance in small data samples and difficulty in positioning big data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162480A_ABST
    Figure CN120162480A_ABST
Patent Text Reader

Abstract

The invention discloses a target object description fusion framework, device, equipment and medium under size data, and relates to the technical field of social simulation, and the target object description fusion framework comprises the following steps: obtaining description data of a target object of a target community based on an active investigation mode, and obtaining probability distribution of each dimension; obtaining description data of a target object of the target community provided by the mobile internet, and obtaining probability distribution of each dimension; acquiring actual probability distribution of a specific dimension as reference probability distribution, determining an optimal weight parameter of each distribution fusion algorithm based on the reference probability distribution, and further determining an optimal distribution fusion algorithm; and fusing the probability distribution of each dimension obtained from the mobile internet with the probability distribution of each dimension obtained in the active investigation mode by using an optimal distribution fusion algorithm and the corresponding optimal weight parameter. According to the method, more comprehensive crowd portrait description can be realized, and social simulation is effectively supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of social simulation technology, and particularly relates to a target object description fusion framework, device, equipment and medium under large and small data. Background Art

[0002] Social simulation is an effective method for solving urban governance problems. By modeling entities such as people, vehicles, roads, buildings, etc., and simulating the interactions between entities, the operation of society can be effectively simulated, social experiments can be carried out, and the formulation of urban governance policies can be explored. In social simulation based on agent modeling, in order to effectively model people, a population portrait is required as a support to construct corresponding agents.

[0003] Currently, the research on population portraits mainly focuses on social media population portraits, that is, analyzing users based on their behaviors such as posting, commenting, liking, following, and forwarding on social media to form population portraits of social media users, and further characterizing aspects such as user consumption preferences and hobbies can be extracted. There is a large amount of related work on user portraits and behavior analysis based on social networks. However, social simulation often focuses on a certain geographical area, and mainly studies the simulation, interaction and emergence phenomena of agents in this area. Social media users are not restricted by geography, making it difficult to locate users in a certain area and then conduct portrait characterization. Moreover, there is an imbalance in the geographical distribution of social media users. Even if users in a certain area can be filtered out, they may not be representative due to the small number, and the population portrait required for social simulation cannot be effectively calculated. Therefore, other methods must be used to obtain the population portrait required for social simulation.

[0004] At present, the methods for obtaining population portraits are mainly divided into two categories. One is to depict based on questionnaire surveys, and the other is to depict based on the mining and analysis of big data on the mobile Internet. Currently, some relevant research scholars have proposed a method for digital portraits of urban populations based on big data (i.e., the big data survey method), which conducts population portraits of urban residents from three dimensions of time, space, and behavior based on spatio-temporal analysis, and analyzes the spatio-temporal behavior characteristics and laws of urban populations. This portrait only analyzes and portrays the three dimensions related to time and space, focuses on its application in urban planning, lacks the portrayal of basic population attributes such as gender, age, education level, consumption, etc., and cannot be directly used for community simulation. There are also relevant research scholars who have proposed using the method of spatio-temporal behavior analysis to conduct population portraits of communities through interview questionnaires. Based on the questionnaire data, four indicators in the three aspects of time, space, and behavior (single-day rhythm, benchmark ground polarization, activity period, main behavior pattern) are defined to depict the community population, and nine types of typical populations are obtained through clustering. The main behavior patterns among them are park rest, community leisure, venue culture and sports, and consumer shopping. The main purpose of this community portrait is to clarify the leisure activity needs of community residents and then give strategic suggestions for urban public facility planning. This portrait also lacks basic population attribute information, the sample data volume is very small and there are corresponding biases, and it cannot be directly used for social simulation.

[0005] The method relying on questionnaire surveys (i.e., the small data survey method) can be directly targeted at the areas of interest, with accurate population positioning, but the collection cost is relatively high, and during weekdays, the respondents are mainly middle-aged and elderly people, so it is not convenient to adopt on a large scale. With the development of the mobile Internet, mobile phones have become an indispensable part of life, generating a large amount of big data on mobile Internet such as positioning, consumption, and entertainment. Mining these big data on the mobile Internet can obtain portraits of the population. Since the data collected by the mobile Internet mainly reflects the information of teenagers, young people, and middle-aged groups using smart phones, there is also a sample imbalance in the data based on the mobile Internet.

[0006] Big data on the Internet and small data from questionnaire surveys complement each other to a certain extent. Integrating the two can achieve a more comprehensive portrait of the population. Some relevant research scholars have proposed a method for community population portraits based on the integration of multi-source data. First, the small data from questionnaire surveys is converted to the same portrayal level as the big data, and then the big and small data are integrated to achieve a more comprehensive portrait of the community population. However, when integrating the portraits of big and small data, what specific integration method to adopt to generate a more comprehensive population portrait has become an urgent problem to be solved currently. Summary of the Invention

[0007] This application provides a fusion framework, device, equipment, and medium for describing target objects under big and small data, which can achieve a more comprehensive portrayal of the population and effectively support social simulation.

[0008] In a first aspect, an embodiment of the present application provides a target object description fusion framework for large and small data. The target object description fusion framework for large and small data includes:

[0009] Based on the electronic map selection, determine the target community for which object description is required, and obtain the description data of the target objects in the target community through an active survey method to obtain the probability distribution of each dimension;

[0010] Call the API interface function to obtain the description data of the target objects in the target community provided by the mobile Internet to obtain the probability distribution of each dimension;

[0011] Obtain the actual probability distribution of a specific dimension as the reference probability distribution, determine the optimal weight parameters of each distribution fusion algorithm based on the reference probability distribution, and then determine the optimal distribution fusion algorithm;

[0012] Use the optimal distribution fusion algorithm and the corresponding optimal weight parameters to fuse the probability distribution of each dimension obtained from the mobile Internet with the probability distribution of each dimension obtained through the active survey method.

[0013] Combined with the first aspect, in an embodiment, the dimensions include gender, age group, education level, age group of children, life stage, whether there is a car, consumption level segment, commuting mode, hotel consumption level segment, hotel type, catering consumption level segment, whether a frequent business traveler, whether traveling abroad, travel month, travel distance, travel destination, food type, shopping place, sports and fitness type, entertainment and leisure type, tourist attraction type, intelligent mobile terminal brand, intelligent mobile terminal price.

[0014] Combined with the first aspect, in an embodiment, the calling the API interface function to obtain the description data of the target objects in the target community provided by the mobile Internet to obtain the probability distribution of each dimension specifically includes:

[0015] Through the electronic map, select and determine the scope of the target community for which object description is required;

[0016] Call the API interface function to obtain the description data of the target objects in the target community provided by the mobile Internet within a unit time to obtain the probability distribution of the target objects in each dimension.

[0017] Combined with the first aspect, in an embodiment,

[0018] The distribution fusion algorithms include a linear fusion algorithm, a log-linear fusion algorithm, and a Holder fusion algorithm;

[0019] The specific calculation method of the linear fusion algorithm is:

[0020]

[0021] Among them, represents the fusion probability distribution of the k-th dimension, θ represents the variable of the k-th dimension, represents the probability distribution of the k-th dimension obtained from the mobile Internet, represents the probability distribution of the k-th dimension obtained through the active investigation method, ω1 and ω2 represent weight parameters, and ω1 + ω2 = 1, ω1 ≥ 0, ω2 ≥ 0;

[0022] The specific calculation method of the log-linear fusion algorithm is as follows:

[0023]

[0024] Among them, C1 represents the normalization constant, and

[0025] The specific calculation method of the Holder fusion algorithm is as follows:

[0026]

[0027] Among them, C2 represents the normalization constant, and α represents the adjustment coefficient, the value range is a real number, and α ≠ 0, ω1 + ω2 = 1, ω1 ≥ 0, ω2 ≥ 0.

[0028] Combined with the first aspect, in an implementation manner, the actual probability distribution of a specific dimension is obtained as the reference probability distribution, and the optimal weight parameters of each distribution fusion algorithm are determined based on the reference probability distribution, and then the optimal distribution fusion algorithm is determined. Among them, for the determination of the optimal weight parameters of the distribution fusion algorithm, it specifically includes:

[0029] Obtain the existing publicly available real description data, and use the actual probability distribution of the specific dimension in the real description data as the reference probability distribution;

[0030] Through the distribution fusion algorithm, the probability distribution of the specific dimension obtained through the active investigation method is fused with the probability distribution of the specific dimension obtained from the mobile Internet to obtain the weighted parameter fusion probability distribution of the specific dimension;

[0031] Calculate the distance between the weighted parameter fusion probability distribution of the specific dimension and the reference probability distribution, and with the goal of minimizing the distance between the two, determine the optimal value of the weight parameter in the current weighted parameter fusion probability distribution of the specific dimension, and use the determined optimal value as the optimal weight parameter of the corresponding distribution fusion algorithm.

[0032] In combination with the first aspect, in one implementation, the actual probability distribution of a specific dimension is obtained as the reference probability distribution, and the optimal weight parameters of each distribution fusion algorithm are determined based on the reference probability distribution, and then the optimal distribution fusion algorithm is determined. Among them, for the determination of the optimal distribution fusion algorithm, it specifically includes:

[0033] The distribution fusion algorithm after determining the optimal weight parameters is used as the determined distribution fusion algorithm;

[0034] Through the determined distribution fusion algorithm, the probability distribution of a specific dimension obtained by the active investigation method is fused with the probability distribution of the specific dimension obtained from the mobile Internet, and the fused probability distribution is calculated;

[0035] The distances between the fused probability distributions corresponding to each determined distribution fusion algorithm and the reference probability distribution are calculated respectively, and the determined distribution fusion algorithm corresponding to the minimum distance is used as the optimal distribution fusion algorithm.

[0036] In combination with the first aspect, in one implementation,

[0037] When calculating the distance between two probability distributions, the JS divergence is used as the measure of the distance;

[0038] The specific expression of the JS divergence is:

[0039]

[0040] where D JS (p‖q) represents the JS divergence between two probability distributions, p represents one of the probability distributions, q represents the other probability distribution, and D KL represents the KL divergence.

[0041] In the second aspect, an object description fusion device for large and small data provided by an embodiment of the present application includes:

[0042] The first acquisition module is used to determine the target community that needs to be described based on the electronic map, and obtain the description data of the target object in the target community through the active investigation method to obtain the probability distribution of each dimension;

[0043] The second acquisition module is used to call the API interface function to obtain the description data of the target object in the target community provided by the mobile Internet to obtain the probability distribution of each dimension;

[0044] The determination module is used to obtain the actual probability distribution of a specific dimension as the reference probability distribution, determine the optimal weight parameters of each distribution fusion algorithm based on the reference probability distribution, and then determine the optimal distribution fusion algorithm;

[0045] A fusion module, which is used to fuse the probability distributions of each dimension obtained from the mobile Internet with the probability distributions of each dimension obtained by the active investigation method by using the optimal distribution fusion algorithm and the corresponding optimal weight parameters.

[0046] In a third aspect, an embodiment of the present application provides a target object description fusion device under large and small data. The target object description fusion device under large and small data includes a processor, a memory, and a target object description fusion program under large and small data stored on the memory and executable by the processor. When the target object description fusion program under large and small data is executed by the processor, the steps of the target object description fusion framework under large and small data described above are implemented.

[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. A target object description fusion program under large and small data is stored on the computer-readable storage medium. When the target object description fusion program under large and small data is executed by a processor, the steps of the target object description fusion framework under large and small data described above are implemented.

[0048] The beneficial effects brought by the technical solutions provided by the embodiments of the present application include:

[0049] By obtaining the actual probability distribution of a specific dimension as the reference probability distribution, determining the optimal weight parameters of each distribution fusion algorithm based on the reference probability distribution, further determining the optimal distribution fusion algorithm among multiple distribution fusion algorithms, and realizing the fusion of the target object descriptions of large data and small data based on the optimal distribution fusion algorithm and the corresponding optimal weight parameters, a more comprehensive population portrait can be depicted, effectively supporting social simulation. Description of the Drawings

[0050] Figure 1 It is a schematic flowchart of the target object description fusion framework under large and small data of the present application;

[0051] Figure 2 It is a schematic diagram of the functional modules of the target object description fusion device under large and small data of the present application;

[0052] Figure 3 It is a schematic hardware structure diagram of the target object description fusion device under large and small data of the present application. Detailed Embodiments

[0053] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.

[0054] To make the purpose, technical solution and advantages of this application clearer, the following will further describe the embodiments of this application in detail in conjunction with the accompanying drawings.

[0055] In a first aspect, the embodiments of this application provide a target object description fusion framework for large and small data. According to the specific situation of the data, the optimal distribution fusion algorithm is selected, and the optimal weight parameter of the optimal distribution fusion algorithm is determined, and a more comprehensive population portrait is fused to effectively solve the problem of population portrait fusion for large and small data and support social simulation.

[0056] In one embodiment, referring to Figure 1 , Figure 1 is a schematic flowchart of the target object description fusion framework for large and small data in this application. As Figure 1 shown, the target object description fusion framework for large and small data includes:

[0057] S1: Based on the electronic map selection, determine the target community for which object description is required, and obtain the description data of the target objects in the target community through an active survey method to obtain the probability distribution of each dimension;

[0058] S2: Call the API (Application Programming Interface) interface function to obtain the description data of the target objects in the target community provided by the mobile Internet to obtain the probability distribution of each dimension;

[0059] S3: Obtain the actual probability distribution of a specific dimension as the reference probability distribution, determine the optimal weight parameter of each distribution fusion algorithm based on the reference probability distribution, and then determine the optimal distribution fusion algorithm;

[0060] S4: Use the optimal distribution fusion algorithm and the corresponding optimal weight parameter to fuse the probability distribution of each dimension obtained from the mobile Internet with the probability distribution of each dimension obtained through the active survey method.

[0061] It should be noted that obtaining the description data of the target objects in the target community through the active survey method refers to small data, and obtaining the description data of the target objects in the target community from the mobile Internet refers to big data.

[0062] In this application, dimensions include gender, age group, education level, age of children, life stage, whether having a car, consumption level segment, commuting mode, hotel consumption level segment, hotel type, catering consumption level segment, whether being a frequent business traveler, whether traveling abroad, month of travel, travel distance, travel destination, type of cuisine, shopping venue, type of sports and fitness, type of entertainment and leisure, type of tourist attraction, brand of intelligent mobile terminal, and price of intelligent mobile terminal. For example, for a certain target object, the obtained descriptive data are: male gender, age within the range of 25 - 30 age group, undergraduate education level, and other dimension information. Of course, for some special cases, such as a target object having no children, in the dimension of age of children, the descriptive data is "no children", and then calculate the probability distribution of the descriptive result of "no children" in the dimension of age of children later.

[0063] When obtaining small data, for the convenience of actual operation, based on the electronic map, the target community area can be selected as a residential community, that is, conduct a questionnaire survey within the scope of the residential community. Conduct a questionnaire survey in a set number of residential communities, collect data related to the population portrait, design a questionnaire including the above 23 - dimension questions, and obtain the descriptive data of the target objects in the target community based on the active survey method. As shown in Table 1 below, it is a partial example of the questionnaire.

[0064] Table 1

[0065] Serial number Gender Age (segmented) Educational background Life stage Age of children 1 Female 31 - 35 years old High school Freelance 2 - 5 years old 2 Male 31 - 35 years old Junior high school Freelance 6 - 12 years old 3 Male 40 - 60 years old Junior high school Freelance Others 4 Female 25 - 30 years old Bachelor's degree Office worker 2 - 5 years old 5 Male Over 61 years old Junior high school Retired Others 6 Male 31 - 35 years old Associate degree Office worker No children 7 Female Over 61 years old Primary school Retired Others 8 Male 40 - 45 years old Bachelor's degree Office worker 6 - 12 years old 9 Female Over 61 years old Primary school Retired Others 10 Female Over 61 years old Primary school Retired Others 11 Female Over 61 years old Primary school Retired Others

[0066] Each row in Table 1 represents the information of a resident. In the dimension of education level, the descriptive results include 7 categories: primary school, junior high school, high school, junior college, undergraduate, master, and doctor. Assume that the sample only has these 11 pieces of data, and introduce the calculation of the probability distribution of small data in specific dimensions. Among the 11 target objects, count the number of target objects in each education - level category respectively, and get 4 people with primary - school education level, 3 people with junior - high school education level, 1 person with high - school education level, 1 person with junior - college education level, 2 people with undergraduate education level, 0 people with master education level, and 0 people with doctor education level. Therefore, the probability distribution of the descriptive results in the education - level dimension calculated from the questionnaire survey sample is: primary school: 4 / 11 = 0.3636, junior high school: 3 / 11 = 0.2727, high school: 1 / 11 = 0.0909, junior college: 1 / 11 = 0.0909, undergraduate: 2 / 11 = 0.1818, master: 0 / 11 = 0.0, doctor: 0 / 11 = 0.0.

[0067] Obviously, Table 1 above is for the case of single - dimension selection. For example, in the gender dimension, a certain target object can only select male or female. When there is a multi - dimension selection case, such as for the commuting mode dimension, a certain target object can select both bus and subway at the same time. For the multi - dimension selection case, when calculating the probability distribution of each dimension, the processing method refers to the Chinese patent with the application number CN202411184960.8 and the title "Target Object Description Method, Device and Equipment Based on Multi - source Data Fusion".

[0068] Further, in one embodiment, an API interface function is called to obtain the description data of the target objects in the target community provided by the mobile Internet, and the probability distribution of each dimension is obtained, specifically including:

[0069] S201: Through the electronic map, select and determine the range of the target community for which object description is to be carried out;

[0070] S202: Call the API interface function to obtain the description data of the target objects in the target community provided by the mobile Internet within a unit time, and obtain the probability distribution of the target objects in each dimension.

[0071] The description data of the target objects in the target community can be obtained through a portal website, an instant messaging operator, etc., according to the relevant publicly available data uploaded or recorded by the target objects in the target community through the mobile Internet.

[0072] Specifically, by delimiting the range of the target community for which object description is to be carried out on the electronic web map, and calling the API function, the statistical distribution of the community population in the above - mentioned 23 dimensions on a daily basis is obtained. For privacy protection, some data are subjected to interval - based fuzzy processing.

[0073] After obtaining the probability distribution of each dimension (big data) from the mobile Internet and the probability distribution of each dimension (small data) through the active investigation method, the next step is to select a candidate distribution fusion algorithm according to the situation of big and small data, and determine the benchmark probability distribution, the optimal distribution fusion algorithm, and the optimal weight parameters of the optimal distribution fusion algorithm.

[0074] For the probability distribution of each dimension obtained from the mobile Internet, it records the probability distribution of the target objects in 23 dimensions Similarly, for the probability distribution of each dimension obtained through the active investigation method, it records the probability distribution of the target objects in 23 dimensions The maximum value of k is 23. For and the θ in, it represents the variable under the corresponding dimension, including nominal and ordinal variables. For example, for It represents the probability distribution of the first dimension obtained from the mobile Internet. If the first dimension is the gender dimension, then θ in takes values among male or female. In practical applications, for the convenience of calculation, the value-taking situations of θ in each dimension can be assigned accordingly. For example, when θ represents male, the value is 1; when θ represents female, the value is 2. For simplicity, the variables in each dimension are all represented by θ.

[0075] When fusing the probability distributions of each dimension obtained from the mobile Internet with the probability distributions of each dimension obtained through the active investigation method, it is necessary to determine an optimal distribution fusion algorithm and the corresponding optimal weight parameters in the optimal distribution fusion algorithm, so as to use the determined optimal distribution fusion algorithm and optimal weight parameters on each dimension to fuse the probability distributions under the large and small data of the current dimension.

[0076] Specifically, determine a distribution fusion algorithm g: (q 1 , q 2 ) → q 3 , where q 1 represents the probability distribution of a certain dimension obtained from the mobile Internet, q 2 represents the probability distribution of a certain dimension obtained through the active investigation method, and q 3 represents the fused probability distribution. After determining the optimal distribution fusion algorithm g and the corresponding optimal weight parameters based on the benchmark dimension and applying them to all dimensions, the fused descriptions of the target objects under all dimensions can be obtained, that is:

[0077]

[0078] Among them, represents the fused probability distribution of the k-th dimension, θ represents the variable of the k-th dimension, represents the probability distribution of the k-th dimension obtained from the mobile Internet, represents the probability distribution of the k-th dimension obtained through the active investigation method. Therefore, the key to probability distribution fusion lies in the determination of the distribution fusion algorithm g and the corresponding weight parameters.

[0079] In this application, the candidate distribution fusion algorithms include the linear fusion algorithm, the log-linear fusion algorithm, and the Holder fusion algorithm;

[0080] The specific calculation method of the linear fusion algorithm is:

[0081]

[0082] Among them, represents the fused probability distribution of the k-th dimension, θ represents the variable of the k-th dimension, denotes the probability distribution of the k-th dimension obtained from the mobile Internet, denotes the probability distribution of the k-th dimension obtained by the active survey method, ω1 and ω2 denote weight parameters, and ω1 + ω2 = 1, ω1 ≥ 0, ω2 ≥ 0; the linear fusion algorithm is to perform weighted summation on the probability distributions to be fused through a set of weight coefficients, specifically for the fusion of large and small data distributions.

[0083] The specific calculation method of the log-linear fusion algorithm is as follows:

[0084]

[0085] where C1 represents the normalization constant, and The log-linear fusion algorithm is to multiply the probability distributions to be fused through a set of weight coefficients by a power function to obtain the fused probability distribution.

[0086] The specific calculation method of the Holder fusion algorithm is as follows:

[0087]

[0088] where C2 represents the normalization constant, and α represents the adjustment coefficient, and its value range is real numbers, and α ≠ 0, ω1 + ω2 = 1, ω1 ≥ 0, ω2 ≥ 0.

[0089] In actual applications, the distribution fusion algorithms used for probability distribution fusion are relatively mature. The key lies in the selection of the distribution fusion algorithm and the determination of the weight parameters therein. Therefore, this application focuses on studying the selection of the optimal distribution fusion algorithm and determining the weight parameters of the optimal distribution fusion algorithm, so as to realize the population portrait fusion of large and small data.

[0090] Further, in one embodiment, the actual probability distribution of a specific dimension is obtained as the benchmark probability distribution, and the optimal weight parameters of each distribution fusion algorithm are determined based on the benchmark probability distribution, and then the optimal distribution fusion algorithm is determined. Among them, for the determination of the optimal weight parameters of the distribution fusion algorithm, it specifically includes:

[0091] S301: Obtain the existing publicly available real description data, and use the actual probability distribution of the specific dimension in the real description data as the benchmark probability distribution, such as the age distribution obtained from the census;

[0092] S302: Through the distribution fusion algorithm, fuse and calculate the probability distribution of the specific dimension obtained by the active survey method with the probability distribution of the specific dimension obtained from the mobile Internet to obtain the weighted parameter fusion probability distribution of the specific dimension;

[0093] S303: Calculate the distance between the weight parameter fusion probability distribution of a specific dimension and the reference probability distribution, determine the optimal value of the weight parameter in the weight parameter fusion probability distribution of the current specific dimension with the goal of minimizing the distance between the two, and use the determined optimal value as the optimal weight parameter of the corresponding distribution fusion algorithm.

[0094] Specifically, we first need to select a benchmark variable, which is a public population portrait variable supported by actual data. For example, we can select an age variable, because the public national census data contains the real overall age data of the national population, which can be used as a benchmark. For example, we obtain the census data of the target community, and use the actual probability distribution of the age dimension in the census data of the province where the target community is located as the benchmark probability distribution. After determining the benchmark probability distribution, we use the linear fusion algorithm, the log-linear fusion algorithm, and the Holder fusion algorithm to fuse the probability distribution of the age dimension of the target community obtained by the active survey with the probability distribution of the age dimension of the target community obtained from the mobile Internet, and obtain a fused probability distribution of a specific dimension (i.e., the age dimension). That is, the linear fusion algorithm calculates a fused probability distribution on a specific dimension, the log-linear fusion algorithm calculates a fused probability distribution on a specific dimension, and the Holder fusion algorithm calculates a fused probability distribution on a specific dimension. Since the weight parameters in the distribution fusion algorithm have not been determined at this time, the calculated fused probability distribution result of the specific dimension also contains weight parameters ω1, ω2, etc.

[0095] Next, the distance between the fusion probability distribution of a specific dimension and the reference probability distribution is calculated, and the optimal value of the weight parameter in the current fusion probability distribution of a specific dimension is determined with the goal of minimizing the distance between the two, and the determined optimal value is used as the optimal weight parameter of the distribution fusion algorithm corresponding to the current fusion probability distribution of a specific dimension. For example, the distance between the fusion probability distribution of a specific dimension corresponding to the linear fusion algorithm and the reference probability distribution is calculated, and the optimal value of the weight parameter in the current fusion probability distribution of a specific dimension is determined with the goal of minimizing the distance between the two, and the determined optimal value is used as the optimal weight parameter of the linear fusion algorithm.

[0096] Furthermore, in one embodiment, the actual probability distribution of a specific dimension is obtained as a reference probability distribution, and the optimal weight parameters of each distribution fusion algorithm are determined based on the reference probability distribution, thereby determining the optimal distribution fusion algorithm, wherein the determination of the optimal distribution fusion algorithm specifically includes:

[0097] S311: using the distribution fusion algorithm after determining the optimal weight parameter as the determined distribution fusion algorithm;

[0098] S312: By determining a distribution fusion algorithm, fuse the probability distribution of a specific dimension obtained through the active investigation method with the probability distribution of the specific dimension obtained from the mobile Internet, and calculate the fused probability distribution;

[0099] S313: Calculate the distances between the fused probability distributions corresponding to each determined distribution fusion algorithm and the reference probability distribution respectively, and take the determined distribution fusion algorithm corresponding to the minimum distance as the optimal distribution fusion algorithm.

[0100] For example, through the linear fusion algorithm after determining the optimal weight parameter, fuse the probability distribution of the age dimension obtained through the active investigation method with the probability distribution of the age dimension obtained from the mobile Internet, and calculate the fused probability distribution, denoted as the first fused probability distribution; through the log-linear fusion algorithm after determining the optimal weight parameter, fuse the probability distribution of the age dimension obtained through the active investigation method with the probability distribution of the age dimension obtained from the mobile Internet, and calculate the fused probability distribution, denoted as the second fused probability distribution; through the Holder fusion algorithm after determining the optimal weight parameter, fuse the probability distribution of the age dimension obtained through the active investigation method with the probability distribution of the age dimension obtained from the mobile Internet, and calculate the fused probability distribution, denoted as the third fused probability distribution. Then calculate the distances between the first fused probability distribution and the reference probability distribution, the second fused probability distribution and the reference probability distribution, and the third fused probability distribution and the reference probability distribution respectively, and take the determined distribution fusion algorithm corresponding to the minimum distance as the optimal distribution fusion algorithm. For example, if the distance between the third fused probability distribution and the reference probability distribution is the smallest, then determine the Holder fusion algorithm as the optimal distribution fusion algorithm. After that, perform the fusion of the probability distributions of other dimensions in the large and small data through the optimal distribution fusion algorithm and its optimal weight parameter.

[0101] Furthermore, when calculating the distance between two probability distributions in this application, the JS divergence is used as the measure of the distance; the specific expression of the JS divergence is:

[0102]

[0103] where, D JS (p||q) represents the JS divergence between two probability distributions, p represents one of the probability distributions, q represents the other probability distribution, and D KL represents the KL divergence. The JS divergence is non-negative, symmetric, and takes values between [0,1], and can be used to measure the difference between probability distributions.

[0104] In view of the fact that the questionnaire survey data (small data) and Internet big data have different degrees of bias (sample imbalance problem), but are complementary to each other to a certain extent, fusing the two can obtain a more comprehensive and accurate population portrait. By selecting and determining the optimal weight parameters and the optimal distribution fusion algorithm, the fusion of the probability distributions of each dimension in the large and small data is realized. Specifically, the population portrait distributions of the big data and small data are obtained, the benchmark variables are determined, the optimal fusion method is selected from the candidate methods based on the benchmark variables, and according to the population portraits of the big and small data, the optimal fusion weight parameters of the selected optimal fusion method are determined, and the fused portrait is calculated to achieve a more complete population portrait than the single-source data, providing support for social simulation.

[0105] The target object description fusion framework under the large and small data of the embodiments of the present application obtains the actual probability distribution of a specific dimension as the benchmark probability distribution, determines the optimal weight parameters of each distribution fusion algorithm based on the benchmark probability distribution, and then determines the optimal distribution fusion algorithm among multiple distribution fusion algorithms. Based on the optimal distribution fusion algorithm and the corresponding optimal weight parameters, the fusion of the target object descriptions of the big data and small data is realized, achieving a more comprehensive description of the population portrait and effectively supporting social simulation.

[0106] In a second aspect, the embodiments of the present application further provide a target object description fusion device under the large and small data.

[0107] In one embodiment, referring to Figure 2 , Figure 2 is a schematic diagram of the functional modules of the target object description fusion device under the large and small data of the present application. As Figure 2 shown, the target object description fusion device under the large and small data includes: a first acquisition module, a second acquisition module, a determination module, and a fusion module.

[0108] The first acquisition module is used to select and determine the target community that needs to be described based on the electronic map, and obtain the description data of the target objects in the target community through the active survey method to obtain the probability distribution of each dimension; the second acquisition module is used to call the API interface function to obtain the description data of the target objects in the target community provided by the mobile Internet to obtain the probability distribution of each dimension; the determination module is used to obtain the actual probability distribution of a specific dimension as the benchmark probability distribution, determine the optimal weight parameters of each distribution fusion algorithm based on the benchmark probability distribution, and then determine the optimal distribution fusion algorithm; the fusion module is used to use the optimal distribution fusion algorithm and the corresponding optimal weight parameters to fuse the probability distribution of each dimension obtained from the mobile Internet with the probability distribution of each dimension obtained through the active survey method.

[0109] In a third aspect, an embodiment of the present application provides a target object description fusion device for large and small data. The target object description fusion device for large and small data can be a device with data processing capabilities such as a personal computer (PC), a laptop computer, or a server.

[0110] Referring to Figure 3 , Figure 3 FIG. is a schematic hardware structure diagram of the target object description fusion device for large and small data involved in the solution of the embodiment of the present application. In the embodiment of the present application, the target object description fusion device for large and small data may include a processor, a memory, a communication interface, and a communication bus.

[0111] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.

[0112] The communication interface includes an input / output (I / O) interface, a physical interface, and a logical interface, etc., which are used to implement the interconnection of components inside the target object description fusion device for large and small data, as well as an interface for implementing the interconnection between the target object description fusion device for large and small data and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, an optical fiber interface, an ATM interface, etc.; the user device can be a display screen, a keyboard, etc.

[0113] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0114] The processor can be a general-purpose processor, and the general-purpose processor can call the target object description fusion program stored in the memory and execute the target object description fusion framework provided by the embodiment of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the target object description fusion program is called can refer to the various embodiments of the target object description fusion framework of the present application for large and small data, which will not be elaborated here.

[0115] Those skilled in the art can understand that Figure 3 the hardware structure shown in Figure 3 does not constitute a limitation on this application. It may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0116] In a fourth aspect, an embodiment of this application further provides a computer-readable storage medium.

[0117] Stored on the computer-readable storage medium of this application is a target object description fusion program under size data. When the target object description fusion program under size data is executed by a processor, the steps of the target object description fusion framework under size data as described above are implemented.

[0118] Among them, the process method implemented when the target object description fusion program under size data is executed can refer to the various embodiments of the target object description fusion framework under size data of this application, which will not be elaborated here.

[0119] The terms "including" and "having" and any variations thereof in the description of the specification, claims, and the above-mentioned drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. The descriptions of terms such as "first", "second", and "third" are used to distinguish different objects, etc., and do not represent a sequential order, nor do they limit that "first", "second", and "third" are of different types.

[0120] In the description of the embodiments of this application, terms such as "exemplary", "for example", or "for instance" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of terms such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.

[0121] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.

[0122] In some processes described in the embodiments of the present application, multiple operations or steps appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0124] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A target object description fusion framework under large and small data, characterized in that: The target object description fusion framework under the large and small data includes: Based on the electronic map, the target community that needs to be described is selected and determined, and the description data of the target object in the target community is obtained based on the active survey method to obtain the probability distribution of each dimension; Calling an API interface function to obtain description data of the target object of the target community provided by the mobile Internet, and obtaining the probability distribution of each dimension; The actual probability distribution of a specific dimension is obtained as a reference probability distribution, and the optimal weight parameters of each distribution fusion algorithm are determined based on the reference probability distribution, thereby determining the optimal distribution fusion algorithm; By using the optimal distribution fusion algorithm and the corresponding optimal weight parameters, the probability distribution of each dimension obtained from the mobile Internet is fused with the probability distribution of each dimension obtained by active survey.

2. The target object description fusion framework under large and small data as claimed in claim 1, characterized in that: The dimensions include gender, age group, education level, children's age group, life stage, whether or not a car is owned, consumption level group, commuting method, hotel consumption level group, hotel type, catering consumption level group, whether or not a frequent business traveler, whether or not traveling abroad, travel month, travel distance, travel destination, food type, shopping place, sports and fitness type, entertainment and leisure type, tourist attraction type, smart mobile terminal brand, and smart mobile terminal price.

3. The target object description fusion framework under large and small data as claimed in claim 1, characterized in that: The calling of the API interface function to obtain the description data of the target object of the target community provided by the mobile Internet and obtain the probability distribution of each dimension specifically includes: Through the electronic map, select and determine the target community scope for object description; Call the API interface function to obtain the description data of the target object in the target community provided by the mobile Internet within a unit time, and obtain the probability distribution of the target object in each dimension.

4. The target object description fusion framework under large and small data as claimed in claim 1, characterized in that: The distributed fusion algorithm includes a linear fusion algorithm, a logarithmic linear fusion algorithm, and a Holder fusion algorithm; The specific calculation method of the linear fusion algorithm is: in, represents the fusion probability distribution of the kth dimension, θ represents the variable of the kth dimension, represents the probability distribution of the kth dimension obtained from the mobile Internet, represents the probability distribution of the kth dimension obtained by active investigation, ω1 and ω2 represent weight parameters, and ω1+ω2=1, ω1≥0, ω2≥0; The specific calculation method of the log-linear fusion algorithm is: Where C1 represents the normalization constant, and The specific calculation method of the Holder fusion algorithm is: Where C2 represents the normalization constant, and α represents the adjustment coefficient, the value range is a real number, and α≠0, ω1+ω2=1, ω1≥0, ω2≥0.

5. The target object description fusion framework under large and small data as claimed in claim 4, characterized in that: The actual probability distribution of a specific dimension is obtained as a reference probability distribution, and the optimal weight parameters of each distribution fusion algorithm are determined based on the reference probability distribution, so as to determine the optimal distribution fusion algorithm. The determination of the optimal weight parameters of the distribution fusion algorithm specifically includes: Obtain existing publicly available real description data, and use the actual probability distribution of a specific dimension in the real description data as the benchmark probability distribution; Through the distribution fusion algorithm, the probability distribution of a specific dimension obtained by active investigation is fused with the probability distribution of a specific dimension obtained from the mobile Internet to obtain a fused probability distribution of a specific dimension with weight parameters. Calculate the distance between the fusion probability distribution containing weight parameters in a specific dimension and the benchmark probability distribution, determine the optimal value of the weight parameter in the fusion probability distribution containing weight parameters in the current specific dimension with the goal of minimizing the distance between the two, and use the determined optimal value as the optimal weight parameter of the corresponding distribution fusion algorithm.

6. The target object description fusion framework under large and small data as claimed in claim 4, characterized in that: The actual probability distribution of a specific dimension is obtained as a reference probability distribution, and the optimal weight parameters of each distribution fusion algorithm are determined based on the reference probability distribution, so as to determine the optimal distribution fusion algorithm. The determination of the optimal distribution fusion algorithm specifically includes: The distribution fusion algorithm after determining the optimal weight parameters is used as the determined distribution fusion algorithm; By determining the distribution fusion algorithm, the probability distribution of the specific dimension obtained by the active survey method is fused with the probability distribution of the specific dimension obtained from the mobile Internet, and the fused probability distribution is calculated; The distance between the fused probability distribution corresponding to each deterministic distribution fusion algorithm and the benchmark probability distribution is calculated respectively, and the deterministic distribution fusion algorithm corresponding to the minimum distance is taken as the optimal distribution fusion algorithm.

7. A target object description fusion framework under large and small data as described in claim 5 or 6, characterized in that: When calculating the distance between two probability distributions, JS divergence is used as the distance measure; The JS divergence is specifically expressed as: Among them, D JS (p‖q) represents the JS divergence between two probability distributions, p represents one of the probability distributions, q represents the other probability distribution, D KL represents the KL divergence.

8. A device for fusion of target object description under large and small data, characterized in that: The target object description fusion device under the large and small data includes: A first acquisition module is used to select and determine a target community for which object description is required based on an electronic map, and obtain description data of target objects in the target community based on an active survey method to obtain probability distribution of each dimension; A second acquisition module is used to call an API interface function to acquire description data of the target object of the target community provided by the mobile Internet, and obtain the probability distribution of each dimension; A determination module is used to obtain the actual probability distribution of a specific dimension as a reference probability distribution, determine the optimal weight parameters of each distribution fusion algorithm based on the reference probability distribution, and then determine the optimal distribution fusion algorithm; The fusion module is used to fuse the probability distribution of each dimension obtained from the mobile Internet with the probability distribution of each dimension obtained by active investigation using the optimal distribution fusion algorithm and the corresponding optimal weight parameters.

9. A device for fusion of target object description under large and small data, characterized in that: The target object description fusion device under large and small data includes a processor, a memory, and a target object description fusion program under large and small data stored in the memory and executable by the processor, wherein when the target object description fusion program under large and small data is executed by the processor, the steps of the target object description fusion framework under large and small data as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a target object description fusion program under large and small data, wherein when the target object description fusion program under large and small data is executed by a processor, the steps of the target object description fusion framework under large and small data as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Target object description method, device and equipment based on multi-source data fusion

    CN119128285A