A user distribution analysis method, apparatus, device, medium and program product

CN122534486APending Publication Date: 2026-08-07CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNITED NETWORK COMM GRP CO LTD
Filing Date
2026-05-12
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]然而,仅依靠信号特征进行用户位置的估算,易出现用户位置偏差大的问题

Benefits of technology

[0008]本发明的有益效果是:通过获取目标区域的网络侧数据,包括最小化路测数据、原始测量报告及信令数据,由于最小化路测数据带有真实的用户的经纬度信息,因此,通过利用最小化路测数据,将最小化路测数据和原始测量报告进行关联,使得原始测量报告对应的用户的位置信息也为准确的用户的位置信息,从而实现用户位置的准确定位。并且,通过位置信息对应的用户标识、信令数据以及目标原始测量报告,可以对用户的位置信息进行分析,以保证用户分布的真实性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122534486A_ABST
    Figure CN122534486A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of user distribution analysis method, device, equipment, medium and program product, communication technical field is involved.The method includes: obtaining the network side data of target area;Wherein, network side data includes: minimization of drive test data, original measurement report and the signaling data corresponding to original measurement report;Based on minimization of drive test data, network side data and original measurement report, the position information of the user corresponding to original measurement report is determined;Based on the position information of user and original measurement report, target original measurement report is determined;Wherein, target original measurement report is original measurement report with position information;Based on the user identifier corresponding to position information, signaling data and target original measurement report, the user distribution corresponding to target area is analyzed.The present application can analyze the position information of user by the user identifier corresponding to position information, signaling data and target original measurement report, to guarantee the authenticity of user distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and specifically to a user distribution analysis method, apparatus, device, medium, and program product. Background Technology

[0002] As the number of mobile communication users continues to expand, accurate analysis of the distribution of users across the network can provide important data support for operators to rationally allocate network resources, plan base station construction, allocate network bandwidth, and formulate emergency rescue measures in the event of natural disasters such as earthquakes and mudslides.

[0003] In related technologies, user location is obtained by combining signal characteristics from network-side user measurement reports with positioning algorithms to achieve user distribution analysis.

[0004] However, relying solely on signal characteristics to estimate user location can easily lead to large deviations in user location.

[0005] Therefore, there is an urgent need for a method that can analyze user location for accurate positioning. Summary of the Invention

[0006] The technical problem to be solved by this invention is that related technologies rely solely on signal characteristics to estimate user location, which is prone to large deviations in user location.

[0007] The technical solution of this invention to solve the above-mentioned technical problems is as follows: a user distribution analysis method, the method comprising: acquiring network-side data of a target area; wherein, the network-side data includes: minimized drive test data, original measurement reports, and signaling data corresponding to the original measurement reports; determining the location information of the user corresponding to the original measurement reports based on the minimized drive test data, network-side data, and original measurement reports; determining the target original measurement reports based on the user location information and the original measurement reports; wherein, the target original measurement reports are original measurement reports with location information; and analyzing the user distribution corresponding to the target area based on the user identifier corresponding to the location information, signaling data, and the target original measurement reports.

[0008] The beneficial effects of this invention are as follows: By acquiring network-side data of the target area, including minimized drive test data, original measurement reports, and signaling data, and since the minimized drive test data contains the actual latitude and longitude information of the users, by using the minimized drive test data and associating it with the original measurement report, the user location information corresponding to the original measurement report is also accurate user location information, thereby achieving accurate user location positioning. Furthermore, by using the user identifier corresponding to the location information, signaling data, and the target original measurement report, the user location information can be analyzed to ensure the authenticity of user distribution.

[0009] Based on the above technical solution, the present invention can be further improved as follows.

[0010] Furthermore, based on the latitude and longitude information in the minimized drive test data and the original measurement report, the location information of the user corresponding to the original measurement report is determined, including: constructing a minimized drive test database based on the first reference signal received power, latitude and longitude information, and minimized sample point data of neighboring cells in the minimized drive test data; and determining the location information from the minimized drive test database based on the second reference signal received power and sample point data of the original measurement report; wherein, the first reference signal received power corresponds one-to-one with the second reference signal received power, and the minimized sample point data of neighboring cells corresponds one-to-one with the sample point data of the original measurement report.

[0011] The beneficial effect of adopting the above-mentioned further scheme is that by using minimized road test data containing latitude and longitude information, constructing a minimized road test database with the first reference signal received power and the minimized sample point data of the neighboring cell, and then matching and associating the second reference signal received power of the original measurement report, the sample point data of the original measurement report, and the minimized road test database, the true location information can be accurately mapped directly to the original measurement report, thereby improving the accuracy of the location results.

[0012] Furthermore, based on the user identifiers corresponding to the location information, signaling data, and the target original measurement report, the user distribution corresponding to the target area is analyzed, including: associating the signaling data with the first target original measurement report through the user identifier; determining the second target original measurement report from the first target original measurement report based on the timestamp information in the signaling data; wherein, the second target original measurement report is the original measurement report in the first target original measurement report whose timestamp is within a preset time range; and analyzing the user distribution corresponding to the target area based on the second target original measurement report.

[0013] The beneficial effect of adopting the above-mentioned further scheme is that by associating signaling data with target original measurement reports containing location information through user identifiers, and filtering out valid measurement reports within a preset time range based on the timestamp in the signaling data, it is possible to achieve accurate filtering and deduplication of user data, ensuring that the samples participating in user distribution analysis are all valid data of the same user within the valid time period, and avoiding duplicate sampling of valid data.

[0014] Furthermore, based on the original measurement report of the second target, the user distribution corresponding to the target area is analyzed, including: grouping the original measurement report of the second target according to user identifiers to obtain multiple groups of original measurement reports; determining the number of sample points corresponding to each group of original measurement reports; based on the number of sample points, determining the third target original measurement report from the multiple groups of original measurement reports whose number of sample points is greater than the number threshold; and based on the third target original measurement report, analyzing the user distribution corresponding to the target area.

[0015] The beneficial effect of adopting the above-mentioned further scheme is that by grouping the original measurement reports of the second target according to user identifiers and selecting valid measurement reports with a sample point count greater than the threshold for participation in user distribution analysis, invalid user data with insufficient sample size and low credibility can be eliminated, thereby further improving the accuracy of user positioning results and distribution statistics results.

[0016] Furthermore, based on the original measurement report of the third target, the user distribution corresponding to the target area is analyzed, including: rasterizing the original measurement report of the third target to obtain the original measurement report of the fourth target; determining the number of users in each grid based on the original measurement report of the fourth target and user identifiers; and analyzing the user distribution corresponding to the target area based on the number of users in each grid.

[0017] The beneficial effect of adopting the above-mentioned further scheme is that by rasterizing the original measurement report of the third target and counting the number of users in each grid with user identifiers, it is possible to transform the discrete and massive user measurement samples into regular and quantifiable regional user distribution data, thereby improving the true user density and distribution characteristics within the target area.

[0018] Furthermore, the method also includes: displaying user distribution on the page as a heatmap; wherein the visual features corresponding to the heatmap are used to characterize the differences in user distribution.

[0019] The beneficial effect of adopting the above-mentioned further solution is to visualize the regional user distribution data in the form of a heat map on the page. By using the different visual features of the heat map, the differences in user density in different regions can be intuitively represented. The quantified raster user distribution data can be transformed into an intuitive and easy-to-understand visualization result, which makes it easier for operators to quickly identify high user density areas and user clustering characteristics.

[0020] Furthermore, the present invention provides a user distribution analysis device, which includes: an acquisition module for acquiring network-side data of a target area; wherein the network-side data includes: minimized drive test data, original measurement reports, and signaling data corresponding to the original measurement reports; a first determination module for determining the location information of the user corresponding to the original measurement report based on the latitude and longitude information in the minimized drive test data, the network-side data, and the original measurement reports; a second determination module for determining a target original measurement report based on the user's location information and the original measurement reports; wherein the target original measurement report is an original measurement report with location information; and an analysis module for analyzing the user distribution corresponding to the target area based on the user identifier corresponding to the location information, the signaling data, and the target original measurement report.

[0021] Furthermore, this application provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the user distribution analysis method described in the first aspect or any corresponding embodiment.

[0022] Furthermore, this application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the user distribution analysis method described in the first aspect or any corresponding embodiment.

[0023] Furthermore, this application provides a computer program product, including computer instructions for causing a computer to execute the user distribution analysis method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0024] Figure 1 A flowchart illustrating the user distribution analysis method provided by this invention; Figure 2 A flowchart illustrating another user distribution analysis method provided by the present invention; Figure 3 A flowchart illustrating another user distribution analysis method provided by the present invention; Figure 4 This is a schematic diagram of the user distribution heatmap provided by the present invention.

[0025] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0026] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0027] By analyzing user distribution, operators can understand user density and activity in different areas, thereby rationally allocating network resources, such as base station construction and bandwidth allocation, to ensure network quality and user experience in high-user-density areas. Secondly, they can identify potential market opportunities and target groups. In other application scenarios, such as natural disasters like earthquakes and mudslides, they can analyze population distribution and effectively formulate emergency rescue measures.

[0028] However, in related technologies, user location is obtained by combining signal characteristics from network-side user measurement reports with positioning algorithms to achieve user distribution analysis.

[0029] However, relying solely on signal characteristics to estimate user location can easily lead to large deviations in user location.

[0030] Based on this, the present invention provides a user distribution analysis method. This method acquires network-side data for a target area, including minimized drive test data, original measurement reports, and signaling data. Since the minimized drive test data contains the actual latitude and longitude information of users, it is used to correlate the minimized drive test data with the original measurement reports. This ensures that the user location information corresponding to the original measurement reports is also accurate, thereby achieving precise user location. Furthermore, by analyzing the user identifiers corresponding to the location information, signaling data, and the target original measurement reports, the user location information can be verified to ensure the authenticity of user distribution.

[0031] According to an embodiment of the present invention, a user distribution analysis method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0032] This embodiment provides a user distribution analysis method that can be used in computer devices such as computers and servers. Figure 1 This is a flowchart of a user distribution analysis method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain network-side data for the target area; wherein, the network-side data includes: minimized drive test data, original measurement report, and signaling data corresponding to the original measurement report.

[0033] The target area can be understood as the geographical range of the user distribution to be analyzed, such as cities, administrative regions, base station coverage areas, etc.

[0034] Network-side data can be understood as operational data automatically collected by base stations and other equipment, requiring no additional reporting from terminals. This network-side data includes: minimal drive test data, raw measurement reports, and the corresponding signaling data for the raw measurement reports.

[0035] Minimized drive test data can be understood as measurement data within the target area. Specifically, minimized drive test data can be 4G data.

[0036] The raw measurement report can be understood as the measurement data periodically reported by the terminal when connected to the network. Specifically, the raw measurement report can be 5G data.

[0037] The signaling data corresponding to the original measurement report can be understood as user service interaction data recorded on the network side. This original measurement report can be used for user identity verification and time validity judgment.

[0038] Specifically, computer equipment can obtain network-side data for the target area from the operator's network platform.

[0039] In one possible implementation, the target area can be District B of City A. Minimum 5 days of drive test data, 3 days of raw measurement reports, and the corresponding signaling data from the raw measurement reports are selected from the target area. The raw measurement reports and their corresponding signaling data must be synchronized.

[0040] Step S102: Based on the latitude and longitude information in the minimized road test data and the original measurement report, determine the location information of the user corresponding to the original measurement report.

[0041] The user's location information corresponding to the original measurement report can be understood as the user's latitude and longitude coordinates. However, the original measurement report does not contain the user's latitude and longitude coordinates; the user's location information corresponding to the original measurement report is determined based on the latitude and longitude information in the minimized road test data.

[0042] As an example, minimized drive test data may include latitude and longitude, the reference signal received power of the primary serving cell, a list of neighboring cells, and the reference signal received power of each neighboring cell. The original measurement report may include the reference signal received power of the primary serving cell, a list of neighboring cells, and the reference signal received power of each neighboring cell. The original measurement report, based on the reference signal received power of the primary serving cell, the list of neighboring cells, and the reference signal received power of each neighboring cell, can accurately locate the user's location information from the minimized drive test data.

[0043] As an example, after determining the minimum road test data and the original measurement report, the Euclidean distance between multiple data points in the original measurement report and the minimum road test data can be determined. The K nearest samples from the multiple data points (such as k=3, k=4, etc.) can be selected, and then a weighted average can be performed based on the K samples from the multiple data points to determine the user's location information.

[0044] Step S103: Based on the user's location information and the original measurement report, determine the target original measurement report; wherein, the target original measurement report is an original measurement report with location information.

[0045] The target's original measurement report is a raw measurement report containing location information. Specifically, after determining the user's location information, this information can be backfilled, so that the original measurement report includes the user's location information, thus obtaining the target's original measurement report.

[0046] In a feasible scenario, 1 million raw measurement reports are obtained. Among these 1 million raw measurement reports, 820,000 raw measurement reports match the user's location information. Therefore, these 820,000 raw measurement reports can be identified as target raw measurement reports.

[0047] Step S104: Based on the user identifiers, signaling data, and original target measurement reports corresponding to the location information, analyze the user distribution corresponding to the target area.

[0048] In this embodiment, signaling data can be used for user identity verification, time validity determination, etc. Specifically, after the user identifier corresponding to the location information, signaling data, and the original target measurement report have all been determined, the user distribution corresponding to the target area can be analyzed based on the user identifier corresponding to the location information, signaling data, and the original target measurement report.

[0049] As an example, the minimized drive test data is cleaned by removing data lacking latitude and longitude or containing abnormal signals. Then, the latitude and longitude, reference signal received power of the primary serving cell, and the reference signal received power of the neighboring cell list are extracted from the minimized drive test data. The reference signal received power of the primary serving cell, the neighboring cell list, and their reference signal received power are also extracted from the original measurement report. The signal features of the two are then matched for similarity, and the latitude and longitude of the drive test data with the highest matching degree is used as the location information of the original measurement report to determine the target original measurement report. Based on the timestamps in the signaling data, data in the target original measurement report whose timestamps fall within a preset time range are selected to analyze the user distribution corresponding to the target area.

[0050] As an example, the original measurement report of the target is segmented according to fixed time windows (such as 5 minutes and 15 minutes); within each fixed time window, it is aggregated by user identifier, with each user counted only once, and abnormal time series data is filtered in combination with signaling timestamps. Then, raster statistics are performed on the valid location data within each fixed time window to analyze the user distribution corresponding to the target area.

[0051] The user distribution analysis method provided by this invention acquires network-side data of a target area, including minimized drive test data, original measurement reports, and signaling data. Since the minimized drive test data contains the actual latitude and longitude information of users, it is used to correlate the minimized drive test data with the original measurement report. This ensures that the user location information corresponding to the original measurement report is also accurate, thereby achieving precise user location. Furthermore, by analyzing the user identifiers corresponding to the location information, signaling data, and the target original measurement report, the user location information can be further analyzed to guarantee the authenticity of the user distribution.

[0052] Based on the latitude and longitude information in the minimized road test data and the original measurement report in step S102 of this embodiment, the location information of the user corresponding to the original measurement report is determined using the following method: Figure 2 The implementation of steps S1021 to S1022 is as follows: Step S1022: Construct a minimized road test database based on the first reference signal received power, latitude and longitude information, and neighboring cell minimized sample point data in the minimized road test data.

[0053] The first reference signal received power can be understood as minimizing the power value of the reference signal received by the terminal from the base station carried in the drive test data.

[0054] Latitude and longitude information can be understood as minimizing the actual geographical location coordinates inherent in the road test data.

[0055] The neighbor cell minimized sample point data can be understood as the relevant measurement data of the neighboring cells around the main serving base station recorded in the minimized drive test data. Specifically, it can include the neighboring cell ID, the neighboring cell reference signal received power, the neighboring cell signal quality, etc.

[0056] The minimized drive test database can be understood as a structured database built on the basis of minimized drive test data, integrating latitude and longitude information, first reference signal received power, and neighboring cell minimized sample point data, after cleaning and standardization. Specifically, the minimized drive test database can be understood as a mapping library between signal features (such as reference signal received power, neighboring cell minimized sample point data, etc.) and location, used for matching and positioning in the original measurement report.

[0057] In this embodiment, under the same geographical location and the same base station coverage, the reference signal power received by the terminal and the neighboring cell minimized sample point data are unique and stable. When the user's location information needs to be determined, a minimized drive test database can be constructed based on the first reference signal received power, latitude and longitude information, and neighboring cell minimized sample point data in the minimized drive test data. This achieves unified processing of the minimized drive test data. After the unified processing of the minimized drive test data, the user's location information can be further determined.

[0058] As an example, acquire minimal drive test data within the target area, ensuring that the minimal drive test data covers the coverage area of ​​all base stations in the target area, and the collection period is no less than 3 days; process the minimal drive test data, removing invalid data, including minimal drive test data without latitude and longitude information, with abnormal reference signal received power (such as greater than -50dBm or less than -120dBm), and with missing neighbor cell sample point data; standardize the format of the cleaned minimal drive test data, where latitude and longitude information can retain 6 decimal places, reference signal received power can retain 1 decimal place, and neighbor cell minimal sample point data can be uniformly recorded as key-value pairs of neighbor cell ID and neighbor cell reference signal received power; then construct the minimal drive test database.

[0059] Step S1023: Based on the second reference signal received power and the sample point data of the original measurement report, determine the location information from the minimized road test database; wherein, the first reference signal received power corresponds one-to-one with the second reference signal received power, and the minimized sample point data of the neighboring cell corresponds one-to-one with the sample point data of the original measurement report.

[0060] Although the original measurement report does not carry latitude and longitude, it contains signal characteristics that correspond one-to-one with the minimized drive test data. For the same user terminal in the same geographical location, the received reference signal power and neighbor cell information will not change significantly. Moreover, the coverage areas of 4G and 5G base stations highly overlap. By matching the signal characteristics of the original measurement report with the signal characteristics of the minimized drive test database, the corresponding latitude and longitude information can be found, and the location backfilling can be completed.

[0061] In one possible implementation, the second reference signal received power and sample point data of each data point in the original measurement report are extracted. Based on the one-to-one correspondence between the first and second reference signal received powers and the one-to-one correspondence between the minimized neighbor cell sample point data and the sample point data of the original measurement report, a matching threshold can be set. This matching threshold can be that the deviation between the first and second reference signal received powers is no greater than 2 dBm, and the overlap between the minimized neighbor cell sample point data and the sample point data of the original measurement report is no less than 80%. The second reference signal received power and sample point data of each original measurement report are input into a minimized drive test database, and minimized drive test data that meet the matching threshold are queried.

[0062] If a unique minimum road test data point that meets the criteria is found, its latitude and longitude are extracted as the location information of the original measurement report. If multiple minimum road test data points that meet the criteria are found, their average latitude and longitude are taken as the final location information. If no minimum road test data point that meets the criteria is found, the original measurement report is determined to be invalid data and is discarded.

[0063] The user distribution analysis method provided in this invention integrates the first reference signal received power, latitude and longitude information, and neighboring cell minimized sample point data from the minimized drive test data to construct a minimized drive test database. This provides a real and accurate benchmark for determining user location information, without relying on uncommercial positioning data or manual drive tests, reducing positioning costs and improving benchmark reliability. Simultaneously, by utilizing the one-to-one correspondence between the first and second reference signal received power, and between the neighboring cell minimized sample point data and the original measurement report sample point data, accurate matching between the original measurement report and the database is achieved. This effectively avoids deviations caused by single signal feature positioning, improving the accuracy and reliability of user location information.

[0064] Based on step S104 of this embodiment, which involves analyzing the user distribution corresponding to the target area based on the original measurement report of the second target, and employing methods such as... Figure 3 The implementation of steps S1041 to S1043 is as follows: Step S1041: Associate the signaling data with the original measurement report of the first target using the user identifier.

[0065] Since the user identifier and location information contained in the original measurement report of the first target lack effective time validity verification, it is impossible to determine whether the original measurement report of the first target is valid data generated by the user within the valid time period. Therefore, signaling data can be associated with the original measurement report of the first target, and the original measurement report of the first target can be filtered by using the user identifier and precise timestamp contained in the signaling data. The specific filtering process is described in step S1042.

[0066] Step S1042: Based on the timestamp information in the signaling data, determine the second target original measurement report from the first target original measurement report; wherein, the second target original measurement report is the original measurement report of the first target whose timestamp is within a preset time range.

[0067] Timestamp information can be understood as the precise time recorded in signaling data for user business interactions or data generation. Specifically, timestamp information is used to determine whether the generation time of the original measurement report is consistent.

[0068] The preset time range can be understood as a pre-defined time range. Specifically, the preset time range can be (m1~m2).

[0069] Using the timestamps in the signaling data associated in step S1041 as a benchmark, a preset time range is set, and valid data in the original measurement report of the first target whose timestamps fall within the preset time range are filtered out to obtain the original measurement report of the second target.

[0070] As an example, the original measurement report for the second target can be determined using the following formula.

[0071] ;in, For the original measurement report of the second target, For timestamp information in signaling data, This is the timestamp of the sample point.

[0072] In one possible implementation scenario, based on the user distribution analysis scenario, a fixed preset time range (such as 5 seconds or 10 seconds) is set, the signaling timestamps in the associated dataset and the timestamps of the original measurement reports of the first target are extracted, the time difference between the two is calculated, and the original measurement reports whose time difference is within the preset time range are retained, which are the original measurement reports of the second target.

[0073] In one possible implementation, the user's service type (such as voice service or data service) can be obtained, and the preset time range can be dynamically adjusted according to the user's service type. For example, voice service has a small amount of data, so a time range of 3 seconds is set; data service has a large amount of data, so a time range of 10 seconds is set.

[0074] Step S1043: Based on the original measurement report of the second target, analyze the user distribution corresponding to the target area.

[0075] Data processing and statistical analysis are performed on the original measurement reports of the second target selected in step S1042 to explore the clustering characteristics and density distribution patterns of users in the target area and obtain accurate user distribution results.

[0076] As an example, the original measurement reports of the second target are grouped by user identifier to obtain multiple groups of original measurement reports; the target area is divided into grids (such as geohash level 8), and the location information of the original measurement reports of the second target is mapped to the corresponding grids to analyze the user distribution corresponding to the target area.

[0077] As an example, the location information of the original measurement report of the second target is deduplicated, that is, only one data point is retained for the same user and the same location. A clustering algorithm is used to cluster the deduplicated location data, grouping user locations that are close to each other into a cluster, and each cluster corresponds to a user gathering area. The number of users in each cluster and the coverage area of ​​the cluster are counted to calculate the user density in order to analyze the user distribution corresponding to the target area.

[0078] The user distribution analysis method provided by this invention associates signaling data with the original measurement report of the first target through user identifiers, and integrates the time validity of signaling data with the location information of the original measurement report of the first target. It does not rely on additional positioning equipment or uncommercial data, which not only ensures the compatibility and relevance of the data source, but also avoids the problem of data association confusion.

[0079] Meanwhile, based on the timestamp information of the signaling data, a preset time range is set to filter out the original measurement reports of the second target within this range, effectively eliminating data that crosses time periods, ensuring that the data participating in the user distribution analysis has temporal consistency and validity, and avoiding invalid data interference that leads to distorted analysis results.

[0080] Finally, the user distribution in the target area is analyzed based on the original measurement report of the second target, which has been double-verified (location valid, time valid), thereby accurately analyzing the user distribution.

[0081] Based on the original measurement report of the second target in step S1043 of this embodiment, the user distribution corresponding to the target area is analyzed, and the solution is implemented as in steps a1 to a4: Step a1: Group the original measurement reports of the second target according to the user identifier to obtain multiple groups of original measurement reports.

[0082] Multiple sets of original measurement reports can be understood as several data groups obtained by grouping the original measurement reports of the second target according to user identifiers. Each data group corresponds to a unique user.

[0083] Specifically, although the original measurement reports of the second target have passed time verification, they are massive discrete data containing location information of multiple users. If distribution analysis is performed directly, the same user will be counted multiple times, and data from different users will be mixed up, leading to distorted analysis results. However, the user identifier is the only core field that can distinguish different users. All original measurement reports of the second target for the same user have completely identical user identifiers, and user identifiers of different users are not repeated. Therefore, grouping by user identifier can integrate all valid location data of the same user together. That is, using the user identifier as the sole grouping criterion, the discrete original measurement reports of the second target are classified, so that all valid measurement reports of the same user are grouped into one group, forming multiple groups of original measurement reports.

[0084] Step a2: Determine the number of sample points corresponding to each set of original measurement reports in the multiple sets of original measurement reports.

[0085] The number of sample points can be used to measure the reliability of user location data. A higher number of sample points for a single user indicates richer valid location data generated within a preset time frame, stronger representativeness and stability of the location information, and more accurate user distribution analysis results based on this data. Conversely, a low number of sample points may indicate invalid data generated by users' momentary network access or occasional stops, failing to accurately reflect the actual distribution of users and leading to distorted results if used for distribution analysis. Therefore, counting the number of original measurement reports for the second target included in each group, i.e., the number of sample points, can quantify the reliability of each user's location data.

[0086] Step a3: Based on the number of sample points, determine the third target original measurement report from multiple sets of original measurement reports where the number of sample points is greater than the number threshold.

[0087] The quantity threshold can be a pre-set threshold. Based on the number of sample points counted in step a2, and combined with the preset quantity threshold, groups with a number of sample points greater than the threshold are selected to obtain the original measurement report of the third target.

[0088] As an example, the quantity threshold can be 10. When the quantity threshold is 10, the original measurement report for the third target can be determined using the following formula.

[0089] ;in, For the original measurement report of the third target, This represents the number of sample points.

[0090] Step a4: Based on the original measurement report of the third target, analyze the user distribution corresponding to the target area.

[0091] After determining the original measurement report for the third target, further analysis of the user distribution corresponding to the target area can be performed.

[0092] As an example, the target area is divided into grids (e.g., Geohash Level 8), and the location information of the original measurement report of the third target is mapped to the corresponding grids; users are grouped by grid ID, and duplicate user identifiers are removed from each grid to count the number of unique users in each grid; based on the number of unique users in each grid, the user density distribution within the target area is determined.

[0093] The user distribution analysis method provided in this invention groups the original measurement reports of the second target according to user identifiers, structuring the discrete data to ensure that valid data of the same user are grouped together, avoiding confusion and duplicate statistics of data from different users; then, by counting the number of sample points in each group of original measurement reports, the reliability of user location data is quantified, and the original measurement reports of the third target with sufficient sample points are selected based on the number threshold, effectively eliminating invalid data with insufficient sample size and low reliability, thereby improving the accuracy of user location information.

[0094] Based on the original measurement report of the third target in step a4 of this embodiment, the user distribution corresponding to the target area is analyzed, and the solution is implemented as in steps a41 to a43: Step a41: Rasterize the original measurement report of the third target to obtain the original measurement report of the fourth target.

[0095] The discrete original measurement report of the third target is rasterized, divided into regular grids, and each location data is mapped to the corresponding grid. A grid ID field is added to obtain the original measurement report of the fourth target.

[0096] As an example, considering the target area range and analysis accuracy requirements, a geohash raster level (e.g., level 8) is set. The latitude and longitude information is extracted from the original measurement report of the third target. Using the geohash encoding algorithm, each latitude and longitude is mapped to a corresponding geohash raster, generating a raster ID. This generated raster ID is then added to the corresponding original measurement report of the third target, retaining the original user identifier, location information, timestamp, and other fields to obtain the original measurement report of the fourth target.

[0097] Step a42, based on the original measurement report of the fourth target and user identifiers, determines the number of users within each grid.

[0098] Using the original measurement report of the fourth objective as the data source, and combining the uniqueness of user identifiers, the user identifiers are grouped by grid ID, and duplicates are counted for each grid to determine the number of unique users in each grid.

[0099] As an example, the original measurement reports of the fourth target are grouped by grid ID using grid ID as the grouping key to ensure that all data of the same grid are grouped together; the user identifiers in each grid are deduplicated, and the number of unique user identifiers in each group is the number of users in that grid.

[0100] Step a43: Analyze the user distribution corresponding to the target area based on the number of users in each grid.

[0101] After determining the number of users in each grid, the user distribution in the target area can be further analyzed.

[0102] As an example, combined Figure 4 As shown, based on the number of users and the area of ​​each grid, the user density of each grid is calculated to determine the degree of user aggregation in different grids; according to the user density, user density levels are divided (e.g., red represents high density, orange represents medium density, blue represents low density, etc.), and level thresholds are set; combined with the geographical boundary of the target area, the distribution pattern of grids of different density levels is analyzed to identify high user density aggregation areas and low user density areas, and user distribution characteristics are summarized.

[0103] As an example, if the original measurement report of the fourth target includes timestamps, the data can be grouped into time-series groups according to preset time intervals (such as 15 minutes), and the number of users in each grid within each time period can be counted. The distribution of grid users and user density in different time periods can be compared to analyze the temporal variation of user distribution. A dynamic heat map can be generated to intuitively present the temporal variation of user distribution and analyze user flow trends.

[0104] The user distribution analysis method provided in this invention, by rasterizing the original measurement report of the third target and counting the number of users in each raster using user identifiers, can transform discrete and massive user measurement samples into regular and quantifiable regional user distribution data, realize the regional aggregation of location data, and solve the problem of scattered location data and difficulty in regional statistics in related technologies.

[0105] Based on this embodiment, the user distribution analysis method further includes: step S104, displaying the user distribution on the page in the form of a heat map; wherein, the visual features corresponding to the heat map are used to characterize the differences in user distribution.

[0106] Heatmaps can be used to visualize user distribution. A heatmap can be understood as a graphic representation of user density distribution within a target area, visually presented through different visual features (such as color intensity and brightness differences).

[0107] Visual features can be understood as the core visual identifiers used by heatmaps to distinguish differences in user distribution, mainly including color depth, brightness, transparency, etc.

[0108] In this embodiment, based on the user distribution data (grid ID, number of users, user density) obtained above, the user distribution quantification data of each grid is transformed into the basic graphic of the heat map through a visualization algorithm.

[0109] As an example, the grid distribution of the target area is accurately mapped to the heat map canvas. The user density of each grid is used as a heat value and assigned to the corresponding pixel unit on the heat map canvas. The user distribution is then displayed on the page as a heat map.

[0110] The user distribution analysis method provided by this invention visualizes regional user distribution data in the form of a heat map on a webpage. It uses different visual features of the heat map to intuitively represent the differences in user density in different regions. It can transform quantified raster user distribution data into intuitive and easy-to-understand visualization results, making it easier for operators to quickly identify high user density areas and user clustering characteristics.

[0111] In some embodiments, the user distribution analysis device of the present invention can be implemented in a combination of hardware and software. As an example, the user distribution analysis device of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the user distribution analysis method of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0112] The modules described in the embodiments of this invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0113] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the user distribution analysis methods described above. That is, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the user distribution analysis method shown in any embodiment of the present invention by calling the computer program.

[0114] In one alternative embodiment, an electronic device is provided, such as Figure 5 As shown, Figure 5 The illustrated electronic device 5000 includes a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, for example, via a bus 5002. Optionally, the electronic device 5000 may further include a transceiver 5004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 5004 is not limited to one type, and the structure of the electronic device 5000 does not constitute a limitation on the embodiments of the present invention.

[0115] Processor 5001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 5001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0116] Bus 5002 may include a path for transmitting information between the aforementioned components. Bus 5002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 5002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus 5002 is represented by only one thick line, but this does not mean that there is only one bus or one type of bus.

[0117] The memory 5003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0118] The memory 5003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 5001. The processor 5001 executes the application code stored in the memory 5003 to implement the content shown in the foregoing method embodiments.

[0119] Among them, electronic devices can also be terminal devices, which can be any device that can install applications, including at least one of smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, smart TVs, and smart in-vehicle devices.

[0120] It should be noted that, Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0121] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned methods.

[0122] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0123] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the user distribution analysis method described above.

[0124] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0125] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0126] The computer-readable storage medium provided in this invention can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EEPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0127] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0128] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0129] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0130] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this invention can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0131] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A user distribution analysis method, characterized in that, The method includes: Acquire network-side data for the target area; wherein, the network-side data includes: minimized drive test data, original measurement reports, and signaling data corresponding to the original measurement reports; Based on the latitude and longitude information in the minimized road test data and the original measurement report, the location information of the user corresponding to the original measurement report is determined; Based on the user's location information and the original measurement report, a target original measurement report is determined; wherein, the target original measurement report is an original measurement report with location information. Based on the user identifier corresponding to the location information, the signaling data, and the original target measurement report, the user distribution corresponding to the target area is analyzed.

2. The user distribution analysis method according to claim 1, characterized in that, The determination of the user's location information corresponding to the original measurement report based on the latitude and longitude information in the minimized road test data and the original measurement report includes: Based on the first reference signal received power, latitude and longitude information, and neighboring cell minimized sample point data in the minimized drive test data, a minimized drive test database is constructed; Based on the second reference signal received power and the sample point data of the original measurement report, the location information is determined from the minimized drive test database; wherein, the first reference signal received power corresponds one-to-one with the second reference signal received power, and the minimized sample point data of the neighboring cell corresponds one-to-one with the sample point data of the original measurement report.

3. The user distribution analysis method according to claim 1, characterized in that, The analysis of user distribution within the target area based on user identifiers corresponding to location information, signaling data, and the original target measurement report includes: The signaling data and the original measurement report of the first target are associated using the user identifier; Based on the timestamp information in the signaling data, a second target original measurement report is determined from the first target original measurement report; wherein, the second target original measurement report is an original measurement report in the first target original measurement report whose timestamp is within a preset time range; Based on the original measurement report of the second target, the user distribution corresponding to the target area is analyzed.

4. The user distribution analysis method according to claim 3, characterized in that, The step of analyzing the user distribution corresponding to the target area based on the original measurement report of the second target includes: The original measurement reports for the second target are grouped according to user identifiers to obtain multiple groups of original measurement reports; Determine the number of sample points corresponding to each set of original measurement reports from multiple sets of original measurement reports; Based on the number of sample points, a third target original measurement report is determined from multiple sets of original measurement reports where the number of sample points is greater than a number threshold. Based on the original measurement report of the third target, the user distribution corresponding to the target area is analyzed.

5. The user distribution analysis method according to claim 4, characterized in that, The analysis of user distribution corresponding to the target area based on the original measurement report of the third target includes: The original measurement report of the third target is rasterized to obtain the original measurement report of the fourth target; Based on the original measurement report of the fourth target and the user identifier, determine the number of users in each grid. Based on the number of users in each grid cell, the user distribution corresponding to the target area is analyzed.

6. The user distribution analysis method according to any one of claims 1-5, characterized in that, The method further includes: The user distribution is displayed on the page as a heatmap; wherein the visual features corresponding to the heatmap are used to characterize the differences in user distribution.

7. A user distribution analysis device, characterized in that, The device includes: The acquisition module is used to acquire network-side data of the target area; wherein, the network-side data includes: minimized drive test data, original measurement report, and signaling data corresponding to the original measurement report; The first determining module is used to determine the location information of the user corresponding to the original measurement report based on the latitude and longitude information in the minimized road test data, the network-side data, and the original measurement report; The second determining module is used to determine a target original measurement report based on the user's location information and the original measurement report; wherein the target original measurement report is an original measurement report with location information. The analysis module is used to analyze the user distribution corresponding to the target area based on the user identifier corresponding to the location information, the signaling data, and the original target measurement report.

8. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the user distribution analysis method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the computer to implement the user distribution analysis method as described in any one of claims 1-6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the user distribution analysis method according to any one of claims 1-6.