Method and system for estimating optimal mixed distribution parameters of vehicle-to-everything communication opportunity intervals

By enumerating hybrid distribution models, using the EM algorithm and KS hypothesis testing combined with Occam's razor, the optimal number of branches and distribution type of the hybrid model for vehicle-to-everything (V2X) communication intervals are determined, solving the problem of unknown hybrid model parameters and improving the model's adaptability and accuracy.

CN116436941BActive Publication Date: 2026-05-01UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2023-03-21
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In the parameter estimation of existing hybrid models of statistical distribution of vehicle-to-everything (V2X) communication intervals, the number of branches of the optimal hybrid model and the statistical distribution type of each branch are unknown, resulting in a lack of unified standards for network protocol design and parameter optimization.

Method used

By enumerating all possible mixed distribution models, the parameters are solved using the expectation-maximization (EM) algorithm. Combined with KS hypothesis testing and Occam's razor, the model with the optimal test statistic and no overfitting is selected, and the optimal number of branches of the mixed model and the statistical distribution type of each branch are determined.

Benefits of technology

The generalization and accuracy of the hybrid distribution model under different time, vehicle density and communication distance scenarios were improved, the model complexity was reduced and the adaptability and accuracy of the model were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116436941B_ABST
    Figure CN116436941B_ABST
Patent Text Reader

Abstract

The application discloses a kind of vehicle networking communication opportunity interval optimal mixed distribution parameter estimation method and system, the method includes: obtaining opportunity interval information from the data set that vehicle latitude and longitude data and its corresponding time data are composed, the arrangement combination of different kinds and number of distribution is carried out, all possible mixed distribution models are enumerated;All distribution models are solved automatically Parameters;Based on K-S hypothesis testing and Occam's Razor principle, select the model that test statistic is optimal and does not overfit from the distribution model that completes parameter solving, that is, give the branch number of optimal mixed model and the statistical distribution type of each branch.The application fully considers the generalization of model under different time, different vehicle density and different communication distance scene, and then proposes a solving method for the multi-objective optimization of reducing model complexity and improving model accuracy, so as to give the branch number of optimal mixed model and the statistical distribution type of each branch.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for estimating the optimal hybrid distribution parameters of opportunity intervals in vehicle-to-everything (V2X) communication. Technical Field

[0001] This invention relates to the field of vehicle networking technology, and in particular to a method and system for estimating the optimal hybrid distribution parameters of vehicle networking communication opportunity intervals. Background Technology

[0002] Vehicle-to-everything (V2X) networks are highly dynamic networks where communication between vehicles relies on communication opportunities that arise when they meet (approach within a certain distance). Based on these communication opportunities, V2X networks can perform hop-by-hop forwarding of data packets, forming a delay-tolerant network (DTN). The distribution model parameters of the communication opportunity intervals in V2X networks are of significant value for network protocol design and parameter optimization. Among these, the hybrid distribution model is currently considered a superior distribution model in academia.

[0003] However, there is a major problem with the parameter estimation of existing hybrid models of statistical distribution of vehicle-to-everything (V2X) communication intervals: the number of branches of the optimal hybrid model and the statistical distribution type of each branch are unknown.

[0004] With the development of economy and technology and the acceleration of urbanization, Intelligent Transportation Systems (ITS) have begun to receive increasing attention in order to better manage vehicles and solve existing problems.

[0005] The constituent entities of ITS (Intelligent Transportation Systems) are mainly divided into two categories: those with vehicle-mounted functions and those without. Examples include intelligent vehicles with vehicle-mounted functions and intelligent infrastructure and base stations without, collectively referred to as "smart cells." ITS utilizes sensing and information communication technologies to collect, process, and transmit relevant road and traffic information through networks, enabling connectivity and communication between vehicles and between vehicles and infrastructure. This breaks down information barriers, facilitating timely dynamic decision-making or real-time dispatching by drivers and traffic managers in appropriate situations. These networks are called Vehicular Ad-hoc Networks (VANETs), derived from Mobile Ad-hoc Networks (MANETs) and applied in the transportation field. Therefore, they inherit the characteristic of MANETs having no central node and possess self-organizing properties.

[0006] MANET requires establishing end-to-end routing before transmitting user data. However, in real-world vehicular ad hoc networks, vehicles are not fixed and may even move at any time, resulting in a lack of a fixed network topology and the inability to guarantee continuous connectivity. This is especially true when dealing with large data volumes, making it difficult to ensure reliable message transmission. Therefore, the traditional model of establishing routes first and then forwarding data cannot be directly applied.

[0007] As a crucial component of ITS (Internet of Vehicles), VANET features typical communication methods such as vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication. The concept of the Internet of Vehicles (IoV) is similar to ITS and VANET, with V2V and V2I being the primary application forms of IoV information communication technologies. The essence of these communications lies in leveraging the encounter opportunities arising from the movement of vehicle nodes to achieve multi-hop information transmission using a "data storage-data carrying-data forwarding" model, without the need for pre-established end-to-end routes. Communication networks with this characteristic are called opportunistic networks. However, this wireless communication is limited by distance N; communication between different vehicles and between vehicles and infrastructure is not constantly established, but primarily occurs when the distance between them is less than the maximum communication distance.

[0008] Inter-contact Time (ICT) refers to the time difference between two adjacent communications between nodes in an opportunistic network, and is an important indicator characterizing communication opportunities in the network. Distribution models of ICT are crucial for routing design, buffer management, and performance evaluation in opportunistic networks. While many scholars have previously studied the distribution of ICT, a unified conclusion regarding its distribution model has yet to be reached.

[0009] The time interval between communication opportunities is influenced by many complex factors, such as node density, maximum wireless communication distance, and road geographic distribution. Current research on the statistical distribution in this area mainly falls into two categories: theoretical and empirical data-based approaches. These approaches summarize models of the time interval between node encounters from different perspectives.

[0010] (1) Theory-based model

[0011] Among the models used to evaluate the performance of mobile ad hoc networks, there are three classic node mobility models: Random Walk Mobility Model (RWM), Random Direction Mobility Model (RDM), and Random Waypoint Mobility Model (RWP).

[0012] RWM is very common in nature, such as Brownian motion. If a node moves to the boundary of the simulated region, it will "bounce" off the simulated boundary like a mirror reflection of light, and then continue moving along this path. RDM is similar to RWM, but the difference is that after reaching the boundary each time, the node will pause for a period of time and then randomly choose a new direction to continue moving. In contrast, the node in the RWP model moves to the next position at a constant speed (the speed and the target position are randomly selected), pauses for a random period of time after reaching the target position, and then repeats the movement process.

[0013] For opportunity networks composed of these three theoretical models, the distribution conclusions obtained are mostly exponential distributions, but there are also cases of power-law distributions: Sharma et al. proposed that the encounter interval follows an exponential distribution based on the RWP model; Cai et al. proved, based on the RWP and RWM models, that in a mobile space considering boundary conditions, the node communication opportunity interval follows an exponential distribution, but if the boundary conditions are not considered, it follows an approximate power-law distribution; Abdulla et al.'s research shows that in the RWP and RDM models in the opportunity network environment, the opportunity interval of mobile nodes can also be approximated as an exponential distribution.

[0014] (2) Model based on actual data

[0015] Many scholars have collected actual data for empirical research, including various datasets such as population, buses, and taxis.

[0016] Zhang et al. pointed out that the time interval between bus encounters based on road granularity showed a periodic structure, which approximately conformed to a mixed normal distribution model; Chaintreau et al. proposed that the time interval between node encounters approximately follows a power law distribution by analyzing crowd movement model data collected by portable devices; In contrast to the results of crowd data, Zhu et al. analyzed the trajectory data of thousands of taxis in Shanghai and found that their time interval between encounters approximately follows an exponential distribution.

[0017] Besides the single distribution mentioned above, many studies support the results of segmented distributions. Li et al. used trajectory data recorded by Beijing taxis equipped with GPS systems to prove that the encounter intervals follow a three-segment distribution, where the middle segment follows a power-law distribution and only the last segment follows an exponential distribution. Other studies, also based on Beijing taxi data, have concluded that there are two-segment distributions: exponential and log-normal. Karaganis et al. found that the chance interval between mobile devices decays according to a power-law within a certain characteristic time, but transforms into exponential decay beyond this characteristic time, and verified this using five types of data including people and vehicles.

[0018] While node motion theory helps simplify problem analysis, these three types of node movement models cannot fully describe the movement of actual nodes. The direction and speed of movement of real nodes, such as vehicles, are not completely random. Therefore, they are not practical for protocol design and performance analysis of real networks. It is more realistic to directly analyze actual data.

[0019] For models built using actual node data, the conclusions of distribution models obtained from different types and scenarios of data also vary. A single distribution is difficult to comprehensively describe the data and has limited responsiveness to dataset features; while segmented distributions exhibit diversity in distribution types, their dividing points require discussion. In contrast, mixed distributions possess the characteristic of "soft boundaries" (i.e., belonging to a certain branch with a certain probability, thus having no clear dividing point), offering advantages in describing complex data and becoming a research hotspot in data modeling across various fields. However, determining the number and types of branches used in a mixed model to ensure good adaptability across various scenarios remains an unresolved issue. Overall, research on time interval distribution models still lacks a unified standard. Summary of the Invention

[0020] This invention provides a method and system for estimating the optimal hybrid distribution parameters of opportunity intervals in vehicle-to-everything (V2X) communication, in order to address the technical problem that current research on time interval distribution models still lacks a unified standard.

[0021] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0022] On one hand, the present invention provides a method for estimating the optimal hybrid distribution parameters of the opportunity interval in vehicle-to-everything (V2X) communication, the method comprising:

[0023] Opportunity interval information is obtained from the dataset consisting of vehicle latitude and longitude data and its corresponding time data. Different types and numbers of distributions are arranged and combined to enumerate all possible mixed distribution models.

[0024] Automated parameter solving for all distribution models;

[0025] Based on KS hypothesis testing and Occam's razor, the model with the optimal test statistic and no overfitting is selected from the distribution models that have completed parameter solving. That is, the number of branches of the optimal mixture model and the statistical distribution type of each branch are given.

[0026] Furthermore, the automated parameter solving for all distribution models includes:

[0027] The Expectation-Maximization (EM) algorithm is used to automatically solve for the parameters of all distribution models.

[0028] Furthermore, the selection of the model with the optimal test statistic and no overfitting from the distribution models that have completed parameter solving, based on KS hypothesis testing and Occam's razor, includes:

[0029] By conducting the KS hypothesis test, the model whose p-value meets the requirements is selected;

[0030] Based on Occam's razor, select the model with the smallest number of branches from the models whose p-values ​​meet the requirements.

[0031] Furthermore, the selection of the model with the optimal test statistic and no overfitting from the distribution models that have completed parameter solving, based on KS hypothesis testing and Occam's razor, includes:

[0032] Define the model effectiveness metric MAC p And the model simplicity index g:

[0033]

[0034] Among them, MAC p (φ) represents the validity index value of model φ; sce represents the scenarios under different conditions, with a total of S scenarios; p value (φ)| sce This represents the maximum p-value that model φ can achieve with optimal parameters Ψ under scenario sce, expressed as:

[0035]

[0036] The value of g represents the number of branches in the hybrid model; the higher the number of branches, the more complex the model.

[0037] Introduce an overall indicator G, as follows:

[0038]

[0039] Where G(φ) represents the overall index value of model φ; g(φ) represents the simplicity index value of model φ;

[0040] Thus, the optimization problem for solving the optimal mixed distribution model is as follows:

[0041] minG(φ)

[0042] stφ∈Φ

[0043] Where Φ represents the set of all possible mixed distribution models;

[0044] By solving the optimization problem, a model with the optimal test statistic and no overfitting is obtained.

[0045] On the other hand, the present invention also provides an estimation system for the optimal hybrid distribution parameters of vehicle-to-everything (V2X) communication opportunity intervals, the estimation system comprising:

[0046] The mixed distribution model building module is used to obtain opportunity interval information from a dataset composed of vehicle latitude and longitude data and its corresponding time data, and to arrange and combine different types and numbers of distributions to enumerate all possible mixed distribution models.

[0047] The model parameter solving module is used to automatically solve the parameters of all distribution models;

[0048] The optimal mixture model selection module is used to select the model with the optimal test statistic and no overfitting from the distribution models that have completed parameter solving, based on KS hypothesis testing and Occam's razor principle. In other words, it gives the number of branches of the optimal mixture model and the statistical distribution type of each branch.

[0049] Furthermore, the model parameter solving module is specifically used for:

[0050] The Expectation-Maximization (EM) algorithm is used to automatically solve for the parameters of all distribution models.

[0051] Furthermore, the optimal hybrid model selection module is specifically used for:

[0052] By conducting the KS hypothesis test, the model whose p-value meets the requirements is selected;

[0053] Based on Occam's razor, select the model with the smallest number of branches from the models whose p-values ​​meet the requirements.

[0054] Furthermore, the optimal hybrid model selection module is specifically used for:

[0055] Define the model effectiveness metric MAC p And the model simplicity index g:

[0056]

[0057] Among them, MAC p (φ) represents the validity index value of model φ; sce represents the scenarios under different conditions, with a total of S scenarios; p value (φ)| sce This represents the maximum p-value that model φ can achieve with optimal parameters Ψ under scenario sce, expressed as:

[0058]

[0059] The value of g represents the number of branches in the hybrid model; the higher the number of branches, the more complex the model.

[0060] Introduce an overall indicator G, as follows:

[0061]

[0062] Where G(φ) represents the overall index value of model φ; g(φ) represents the simplicity index value of model φ;

[0063] Thus, the optimization problem for solving the optimal mixed distribution model is as follows:

[0064] minG(φ)

[0065] stφ∈Φ

[0066] Where Φ represents the set of all possible mixed distribution models;

[0067] By solving the optimization problem, a model with the optimal test statistic and no overfitting is obtained.

[0068] In another aspect, the present invention also provides an electronic device comprising a processor and a memory; wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described method.

[0069] In another aspect, the present invention also provides a computer-readable storage medium storing at least one instruction that is loaded and executed by a processor to implement the above-described method.

[0070] Mixed distribution models offer the greatest advantage in representing complex data compared to other single distribution or piecewise distribution models. However, the determination of the number of branches and the selection of model types in mixed models have not been systematically studied. Therefore, this invention focuses on the statistical analysis of mixed distribution models, considering multiple types of mixed distributions, to explore a method for rapidly estimating the parameters of the optimal mixed distribution model. The beneficial effects include at least the following:

[0071] First, this invention proposes a method for estimating the parameters of a mixed distribution model for chance-interval data. This method enumerates all possible mixed distribution models by permuting and combining different types and numbers of distributions, and then uses the Expectation Maximization Algorithm (EM) to automatically solve for the parameters of all distribution models. A macro p-value index based on the Kolmogorov-Smirnov test is proposed, fully considering the model's generalization ability under different time periods, vehicle densities, and communication distances. Furthermore, a solution method is proposed for multi-objective optimization aimed at reducing model complexity and improving model accuracy, thereby selecting the model with the optimal test statistic and without overfitting, i.e., providing the optimal number of branches and the statistical distribution type of each branch of the mixed model.

[0072] Second, this invention uses a real taxi dataset to empirically verify the method, considering four distribution types previously observed in research: normal distribution, log-normal distribution, exponential distribution, and power-law distribution. Ultimately, two mixed distribution models are obtained: a 3-branch log-normal mixture model and a 2-log-normal plus 1-exponential mixture model. These models demonstrate simplicity and effectiveness across different time, distance, and vehicle density scales. This proves the effectiveness of the technical solution of this invention. Attached Figure Description

[0073] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0074] Figure 1 is a schematic diagram of the execution flow of the method for estimating the optimal hybrid distribution parameters of vehicle-to-everything (V2X) communication opportunity intervals provided in an embodiment of the present invention;

[0075] Figure 2 is a schematic diagram of the macro p-values ​​of each hybrid model provided in the embodiments of the present invention. Where (a) is the macro p value of the hybrid model under the default view, (b) is the EL view of (a) and (c) is the EN view of (a).

[0076] Figure 3 is a box plot of the p-value provided in an embodiment of the present invention.

[0077] Figure 4 is a line graph of the p-value of the hybrid model with the best performance of each branch provided in the embodiment of the present invention; where (a) represents the same vehicle density on different dates, and (b) represents different vehicle densities on the same date.

[0078] Figure 5 is a schematic diagram of the complementary cumulative distribution function of the seven models provided in the embodiment of the present invention (g = 3, d = 150m, M = 300). Detailed Implementation

[0079] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0080] First Embodiment

[0081] This embodiment provides a method for estimating the optimal hybrid distribution parameters of opportunity intervals in vehicle-to-everything (V2X) communication. This method can be implemented by an electronic device, and its execution flow is shown in Figure 1. A massive dataset is input for parameter estimation, ultimately yielding the optimal number of branches and distribution type of the hybrid model. The method includes the following steps:

[0082] S1: Obtain opportunity interval information from the dataset consisting of vehicle latitude and longitude data and its corresponding time data; arrange and combine the distributions of different types and numbers; and enumerate all possible mixed distribution models.

[0083] S2 automatically solves the parameters for all distribution models;

[0084] S3, based on KS hypothesis testing and Occam's razor, selects the model with the optimal test statistic and no overfitting from the distribution models that have completed parameter solving, that is, gives the number of branches of the optimal mixture model and the statistical distribution type of each branch.

[0085] The estimation method for the optimal hybrid distribution parameters of the opportunity interval in vehicle-to-everything (V2X) communication is described in detail below.

[0086] 1. Construct a hybrid distributed model pool

[0087] To find the optimal model, we first need to enumerate all possible scenarios, and this set of all possible hybrid models is called the "hybrid model pool." Subsequent steps simply involve taking hybrid models from the pool and conducting experiments. The constraint parameter required to construct the model pool is the maximum number of branches, g. max And the number of model types K. Two construction methods are described below.

[0088] The first method uses the "ball-throwing problem" from permutation and combination techniques. It assumes there are g balls to throw into k boxes, and the boxes can be empty at the end. The total number of balls is... There are several different ways to throw the ball. This uses the "partition idea," which is equivalent to using K-1 partitions to separate the g balls, but allowing the partitions to be adjacent. That is, select K-1 positions from g+K-1 positions and place K-1 partitions. Returning to the construction of the model pool, the number of branches g in the hybrid model satisfies g∈{1,...,g...} max Therefore, there are a total of} in the pool Different hybrid models.

[0089] Considering the existing conclusions in the current research, the most frequently occurring distributions are normal, log-normal, exponential, and power-law. Therefore, this paper considers these four distributions, with K=4. Let the N-dimensional dataset X follow a g... n Branching normal distribution, g l Branching log-normal distribution, g e Branching exponential distribution, g p A mixed distribution model of a branched power-law distribution, where g n ,g l ,g e ,g p ∈{0,1,…,g max For simplicity, we will denote this hybrid model as:

[0090] In the second construction method, consider a K-bit (g) max +1) base numbers, use brute force to divide them from "0001" to "g". max The values ​​of "000" are enumerated sequentially (taking K=4 as an example), and the sum of the four digits is less than g. max The numbers obtained correspond to g in order of digits from left to right. n g l g e and g p The value of .

[0091] 2. Parameter estimation of mixed models

[0092] After obtaining the target model pool, the mixture models can be extracted sequentially for parameter estimation. This step mainly focuses on the EM algorithm, explaining the implementation principle of parameter estimation.

[0093] Let the N-dimensional ICT dataset X = {x1, x2, ..., x} N The hybrid model is composed of opportunity interval data for each vehicle, meaning that the opportunity interval data for each vehicle is a subset of its subset. The weights of each branch of the hybrid model are denoted as π1, π2, ..., π. gand satisfy Then the observed data x j The probability density function (PDF) is expressed as follows:

[0094]

[0095] Among them, f n f l f e and f p Let represent the PDFs of the normal, log-normal, exponential, and power-law distributions, respectively, and each is defined by a parameter. λ i and a i It is uniquely determined. Therefore, the unknown parameter vector Ψ has a specific form:

[0096]

[0097] The EM algorithm introduces latent parameters, typically the number of mixture models. Through two iterative steps—calculating the expectation (E-step) and the maximum (M-step)—it obtains convergent parameter results. It is primarily used for maximum likelihood estimation and maximum a posteriori estimation of missing data, and is also widely used in cluster analysis in machine learning. For mixture model parameter estimation, the original branch to which the sample data belonged is a latent variable; that is, dataset X does not include all information. Therefore, we introduce a latent variable Z, satisfying the following condition:

[0098]

[0099] Let R(Z) be a certain distribution of variable Z, then we have R(z) ij )≥0, R(Z) can be obtained from observed data x j Perform the calculation: Using Jensen's inequality, the derivation of the log-maximum likelihood function for the unknown parameters is as follows:

[0100]

[0101] We denote the lower bound of equation (4) as a new function Q as an auxiliary function for solving the problem. Then, we only need to update the maximum value of Q to continuously update the lower bound of the likelihood function, thereby approximating the optimal parameters. The expression for the Q function is shown in equation (5).

[0102]

[0103] In summary, the main steps of the EM algorithm for solving the parameters of a mixture model can be summarized as follows:

[0104] 1) Initialize parameter Ψ (0)

[0105] 2) Step E: Calculate the auxiliary function Q(Ψ; Ψ) (t) )

[0106] 3) M step: Update parameter Ψ (t+1) This allows the updated Q function to reach its maximum value, i.e., satisfying Q(Ψ). (t+1) Ψ (t) ) = maxQ(Ψ;Ψ (t) )

[0107] 4) Repeat the E-step and M-step until the parameters converge, i.e., ||Ψ (t+1) -Ψ (t) ||Small enough.

[0108] It should be noted here that the following two alternative solutions can also be used for model parameter estimation.

[0109] The first type directly uses a mixture of normal distributions for modeling and parameter estimation, meaning that the statistical distribution type of each branch of the mixture model only considers the normal distribution. Theoretically, by superimposing several normal distribution curves in a certain proportion, any curve can be approximated infinitely. However, for complex real-world data, this approach might require a large number of branches to achieve sufficient accuracy, making the model too complex and difficult to implement.

[0110] The second type involves directly defining the form of the mixed distribution (given the distribution type and number of branches) for parameter estimation. However, the given distribution form may not be suitable for the actual dataset, so the model built using this method is difficult to accurately reflect the data characteristics, and its effectiveness cannot be guaranteed.

[0111] In contrast, the parameter estimation method used in this embodiment can flexibly consider the selection of multiple distribution types and their number of branches, while also imposing restrictions on model complexity, making the optimal mixture model accurate and efficient. It is an effective method to reduce model complexity and improve model accuracy.

[0112] 3. Evaluation of the hybrid model

[0113] After obtaining numerous models and their fitting parameters, it is necessary to determine which to accept or reject based on certain criteria, selecting the relatively best model from the candidate models. To this end, this embodiment uses KS hypothesis testing and Occam's razor to evaluate the model's effectiveness and simplicity, respectively.

[0114] Commonly used model selection methods include the NLL test statistic, the AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) criteria, and the KS (Kolmogorov-Smirnov) test. These methods measure the closeness between theoretical and empirical (actual) distributions by constructing test statistics. The NLL test statistic is the simplest, defined as the negative of the maximum likelihood function; a smaller value indicates a better fit. The AIC statistic is defined as AIC = 2k - 2l, where k is the number of parameters (equivalent to a penalty term), and l is the log-likelihood function value; a smaller value indicates a better fit, but for large sample data, l may be too large, weakening the influence of k. Compared to AIC, the BIC considers the influence of sample size, defined as BIC = klnn - 2l, where n is the sample size. The KS test is a method that directly utilizes data for analysis without requiring parameters; it is based on the Cumulative Distribution Function (CDF) and therefore makes good use of the original data. The KS test statistic is defined as D = max|FF * | represents the maximum difference between the cumulative distribution function of the sample data and the cumulative distribution function of the theoretical data, F and F * The CDF values ​​are for empirical data and theoretical data, respectively. The KS test is based on statistical assumptions; if a low-probability event occurs, there is reason to doubt the validity of the hypothesis and thus reject it. British statistician Ronald Fisher used 1 / 20 as the low-probability standard, and 0.05 is the p-value of the KS test. A p-value greater than 0.05 indicates that the hypothesis cannot be rejected. Since the KS test is a non-parametric test method, this embodiment uses the p-value of the KS test as one of the criteria for judging the reasonableness of the model.

[0115] Occam's Razor, meaning "simple yet effective principle," aims to find the simplest, most computationally efficient, and easiest-to-implement model by avoiding unnecessary entity addition. Theoretically, the more branches a mixture model has, the closer the p-value of the KS test is to 1, and the better the parameter fit. However, this also increases computational complexity and may lead to overfitting.

[0116] This embodiment comprehensively quantifies the above two evaluation indicators, and the specific implementation process is as follows:

[0117] Given an N-dimensional dataset X, we construct a mixture model and estimate its parameters. We limit the maximum number of branches in the mixture model to g. max The mixture model has K types of models, and each distribution contains g.k Branches (g) k (It can be set to 0). Therefore, the total number of branches in the hybrid model is... And g∈{1,...,g max Let Φ denote the set of all possible mixture models, and φ denote one of the possible mixture models. We aim to find one or more optimal mixture models within Φ that are both effective (accurately reflecting the data characteristics) and simple (easy to implement without being overly complex). The former is determined by the p-value in the KS test, while the latter is determined by selecting the model with the fewest branches satisfying the required p-value according to Occam's razor. Therefore, we have selected two parameters as indicators of the effectiveness and simplicity of the mixture model: MAC... p and g. MAC p The macro p-value representing various scenarios is defined as follows:

[0118]

[0119] Here, sce represents scenarios under different conditions, and there are a total of S types. p value (φ)| sce This means that in the SCE scenario, the model φ can achieve the maximum p-value with the optimal parameters Ψ, expressed as:

[0120]

[0121] Equation (7) is essentially the process of using the EM algorithm to solve for the optimal parameters. Looking back at equation (6), the macro p-value is the p-value of the model in all possible scenarios, which can characterize the accuracy of the hybrid model φ under the premise of considering generalization. As for the index g, it is obviously the number of branches of the hybrid model. The higher the number of branches, the more complex the model.

[0122] In real-world scenarios, a model cannot simultaneously be both effective and simple. Therefore, we introduce an overall metric G, as follows:

[0123]

[0124] Thus, we have derived the optimization problem for finding the optimal hybrid model:

[0125]

[0126] The entire framework can be represented in pseudocode as follows:

[0127] Input: Limit the maximum number of branches in the model, g maxThe minimum value of candidate model types K and G is initially set to c = 10. 6

[0128] Output: Optimal mixture model φ optimal

[0129]

[0130] The method of this embodiment will now be verified using a real dataset.

[0131] This embodiment uses a dataset derived from the latitude and longitude data and corresponding time data of 28,590 taxis in a city over seven consecutive days (Monday to Sunday) in May 2010, with a GPS sampling interval of 10 seconds. For simplified calculation, several hundred taxis (including 100, 200, and 300 taxis) were randomly selected from the daily data to calculate the ICT data. Using existing opportunity interval calculation methods, a double-selection and linear interpolation process was employed to obtain opportunity interval information from the original latitude, longitude, and time datasets.

[0132] To ensure the model's generalization ability across different scenarios, we used multiple ICT datasets with varying spatiotemporal scales. The dataset with a total of M vehicles, a maximum communication distance of d, and a date specified as "date" is denoted as "dateMd". The fourteen datasets used in this experiment are listed in Table 1.

[0133] Table 1 Different parameters of the ICT dataset

[0134] Date M (cars) d (m) Abbreviation Monday 300-100m Monday 300-100m Tuesday 300-100m Tuesday 300-100m Wednesday 300-100m Wednesday 300-100m Thursday 300-100m Thursday 300-100m Friday 300-100m Friday 300-100m Saturday 300-100m Saturday 300-100m Sunday 300-100m Sunday 00c-100m Monday 300-150m Tuesday 300-150m Wednesday 300-150m Thursday 300-150m Thursday 300-150m Friday 300-150m Saturday 300-150m Sunday 300-150m surface

[0135] Set the maximum number of branches (g) max =5, the branch model categories include the four types mentioned above: normal distribution (N), log-normal distribution (L), exponential distribution (E), and power-law distribution (P) (K=4). Therefore, the model pool contains a total of Different mixture models were used. For each mixture model, the macro p-value for the dataset in Table 1 was calculated sequentially, and the results are shown in Figure 2. The three-dimensional coordinate axes represent the three distributions N, L, and E, and P is represented by the size of the points in the figure. The color of the points represents the p-value.

[0136] As can be seen from Figure 2(a), all points fall within a triangular pyramid, with p-values ​​ranging from 0 to 0.822. Generally, the larger the number of branches in the model, the larger the p-value. To more clearly illustrate the impact of these four distribution numbers on the overall model performance, Figure 2(b) and Figure 2(c) show the perspectives of Figure 2(a) with the EL and EN planes as the main views, respectively. Figure 2(b) shows that increasing the number of branches in E and L improves the fitting effect of the corresponding mixture model, especially when L increases, the improvement is more significant. Figure 2(c) shows that increasing N does not have a significant impact on the mixture model and may even decrease the p-value. In contrast, the performance of P is unstable; increasing the number of branches in P results in varying performance of the mixture model without a clear pattern.

[0137] To identify the best-performing models for each number of branches, we selected the top three mixture models with the highest p-values ​​when g ranges from 1 to 5. We then plotted box plots of the p-values ​​of these 15 mixture models across different datasets, as shown in Figure 3. It can be seen that the performance difference between models with d values ​​of 100m and 150m is minimal. When g = 1, the p-values ​​are all 0, indicating an insufficient number of branches and ineffective model fitting. When g = 2, only a few models pass the test on some datasets, and the overall fit is unsatisfactory. When g = 3, the macro p-values ​​of L3 and L2E1 both reach 0.6, and they fit well across all datasets, significantly outperforming the third-ranked N1L2. When the number of branches reaches 4 and 5, the p-values ​​are larger, and the box plots are more concentrated, indicating a better model fit. However, the increase in p-values ​​begins to slow down. This suggests that as the number of branches g continues to increase, the complexity of the model increases, but the p-value does not grow linearly. Considering system overhead, further increasing the number of branches is not "cost-effective."

[0138] To explore the optimal number of branches, we extracted the models with the best p-values ​​under each branch and plotted the changes in p-values ​​as line graphs. Figure 4(a) shows the trend of the optimal p-value for the seven datasets in Table 1 under different numbers of branches when d = 100m. As can be seen from the figure, the trends of the seven lines are very similar, and there is a clear "inflection point" at g = 3. When g < 3, the p-value increases significantly with the increase of g, but when g > 3, the increase in p-value slows down or even flattens out. This indicates that when the number of branches is relatively small, g is the main factor affecting the fitting effect of the mixture model; after the number of branches exceeds a certain value, even if g continues to increase, the resulting fitting bonus is limited, and it will increase the complexity of the system implementation, which may be detrimental to practical calculation and application. In addition, vehicle density is also a major factor that may affect the distribution of ICT. We selected vehicles on Tuesday and extracted three other datasets with 100, 200 and 300 vehicles respectively for experiments. The modeling results are shown in Figure 4(b). As can be seen, the inflection point in Figure 4(b) is also at g=3, consistent with Figure 4(a). Furthermore, the p-values ​​at the inflection points are all above 0.6 or higher than 0.8, passing the KS test. Therefore, considering both the fitting effect and system overhead, this invention considers g=3 as the optimal number of branches for the hybrid model.

[0139] We list the top five p-value models with 3 branches on the daily dataset in Table 2 to explore the optimal model. Seven hybrid models appeared more than twice in Table 2: N0L3E0P0, N0L2E1P0, N1L2E0P0, N0L2E0P1, N0L1E1P1, N0L1E2P0, and N1L1E1P0. N0L3E0P0 and N0L2E1P0 almost always occupy the top two positions. Although the other models performed reasonably well, their p-values ​​fluctuated significantly, ranging from 0.25 to 0.66.

[0140] Table 2 shows the top five models and their corresponding p-values ​​(g=3, d=100m, M=300).

[0141]

[0142]

[0143] The following focuses on these seven mixture models. Using the dataset with d = 150m, and keeping other conditions such as the number of vehicles constant, we obtain the data in Table 3. Taking Tuesday's dataset tue300c-150m as an example, we plotted the CCDF (Complementary Cumulative Distribution Function) for each model, as shown in Figure 5. Figure 5 shows that the CCDF curves of each model are very close to the empirical data. Combining the results in Table 3, some models failed the test (p-value less than 0.05), while N0L3E0P0 and N0L2E1P0 remain the two best-performing mixture models.

[0144] Table 3 shows the p-values ​​(g=3, d=150m, M=300) of the seven models and their corresponding datasets.

[0145]

[0146]

[0147] In summary, L3 and L2E1 are the optimal models, with macro p-values ​​of 0.686 and 0.654, respectively.

[0148] Experimental results based on empirical data show that the optimal number of branches for the mixture model is 3. Fewer than 3 branches result in poor fitting, while higher than 3 branches lead to more significant system overhead than improvement in fitting. The optimal mixture models are L3 and L2E1, with macro p-values ​​of 0.686 and 0.654, respectively. Both optimal mixture models contain a large proportion of log-normal distributions, which are heavy-tailed. This may be related to the large values ​​in the ICT data, meaning that some vehicles may not encounter the next vehicle for a long period. The combination of log-normal and exponential distributions is similar to the exponential-log-normal piecewise distributions already discussed in the literature. The fitting effect of combinations of log-normal with normal or power-law distributions is unsatisfactory.

[0149] In summary, this embodiment proposes a method for constructing the optimal hybrid distribution parameters for opportunity intervals in vehicle-to-everything (V2X) networks. It fully considers the model's generalization ability under different time periods, vehicle densities, and communication distances, proposes a macro-p-value index for the model, and then presents a solution method for multi-objective optimization aimed at reducing model complexity and improving model accuracy. This yields the optimal number of branches in the hybrid model and the statistical distribution type of each branch. This method has broad application prospects in the field of V2X networks.

[0150] Second Embodiment

[0151] This embodiment provides an estimation system for the optimal hybrid distribution parameter of opportunity interval in vehicle-to-everything (V2X) communication. This estimation system includes the following modules:

[0152] The mixed distribution model building module is used to obtain opportunity interval information from a dataset composed of vehicle latitude and longitude data and its corresponding time data, and to arrange and combine different types and numbers of distributions to enumerate all possible mixed distribution models.

[0153] The model parameter solving module is used to automatically solve the parameters of all distribution models;

[0154] The optimal mixture model selection module is used to select the model with the optimal test statistic and no overfitting from the distribution models that have completed parameter solving, based on KS hypothesis testing and Occam's razor principle. In other words, it gives the number of branches of the optimal mixture model and the statistical distribution type of each branch.

[0155] The estimation system for the optimal hybrid distribution parameter of vehicle-to-everything (V2X) communication opportunity interval in this embodiment corresponds to the estimation method for the optimal hybrid distribution parameter of V2X communication opportunity interval in the first embodiment described above. The functions implemented by each functional module in the estimation system for the optimal hybrid distribution parameter of V2X communication opportunity interval in this embodiment correspond one-to-one with the process steps in the estimation method for the optimal hybrid distribution parameter of V2X communication opportunity interval in the first embodiment described above. Therefore, it will not be described again here.

[0156] Third Embodiment

[0157] This embodiment provides an electronic device, which includes a processor and a memory; wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment.

[0158] The electronic device can vary considerably depending on its configuration or performance, and may include one or more processors (central processing units, CPUs) and one or more memories, wherein the memories store at least one instruction that is loaded by the processor and executed in accordance with the above method.

[0159] Fourth embodiment

[0160] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc. The instruction stored therein can be loaded and executed by a processor in a terminal.

[0161] Furthermore, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0162] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0163] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal equipment to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams, whereby the instructions that execute on the computer or other programmable terminal equipment provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0164] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0165] Finally, it should be noted that the above description represents a preferred embodiment of the present invention. It should be pointed out that although preferred embodiments have been described, those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles described herein. These improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A method for estimating the optimal hybrid distribution parameters of opportunity intervals in vehicle-to-everything (V2X) communication, characterized in that, The rapid estimation method for the optimal hybrid distribution parameters of the opportunity interval in vehicle-to-everything (V2X) communication includes: obtaining opportunity interval information from a dataset composed of vehicle latitude and longitude data and their corresponding time data; permuting and combining different types and numbers of distributions to enumerate all possible hybrid distribution models; automatically solving the parameters of all distribution models; and selecting the model with the optimal test statistic and no overfitting from the distribution models with completed parameter solutions based on KS hypothesis testing and Occam's razor principle, i.e., providing the number of branches and the statistical distribution type of each branch of the optimal hybrid model. The automatic parameter solution for all distribution models includes: using the expectation-maximization algorithm (EM) to automatically solve the parameters of all distribution models. The selection of the model with the optimal test statistic and no overfitting from the distribution models with completed parameter solutions based on KS hypothesis testing and Occam's razor principle includes: selecting the model whose p-value meets the requirements through KS hypothesis testing; and selecting the model with the smallest number of branches from the models whose p-value meets the requirements according to Occam's razor principle.

2. The method for estimating the optimal hybrid distribution parameters of vehicle-to-everything (V2X) communication opportunity intervals as described in claim 1, characterized in that, The method, based on KS hypothesis testing and Occam's razor, selects the model with the optimal test statistic and no overfitting from the distribution models that have completed parameter solving. This includes defining the model effectiveness index MAC. p And the model simplicity index g: in, Representation Model The effectiveness index value; sce represents the scenario under different conditions, and there are a total of S scenarios; This indicates that in the scenario SCE, the model In optimal parameters The maximum p-value that can be obtained is expressed as: The value of g represents the number of branches in the mixture model; the higher the number of branches, the more complex the model. An overall metric G is introduced as follows: in, Representation Model The overall index value; Representation Model The simplicity index value; thus, the optimization problem for solving the optimal mixed distribution model is as follows: in, This represents the set of all possible mixed distribution models; by solving the optimization problem, a model with the optimal test statistic and without overfitting is obtained.

3. A system for estimating the optimal hybrid distribution parameters of opportunity intervals in vehicle-to-everything (V2X) communication, characterized in that, The rapid estimation system for the optimal hybrid distribution parameters of the opportunity interval in vehicle-to-everything (V2X) communication includes: a hybrid distribution model construction module, used to obtain opportunity interval information from a dataset composed of vehicle latitude and longitude data and their corresponding time data, and to arrange and combine different types and numbers of distributions to enumerate all possible hybrid distribution models; a model parameter solving module, used to automatically solve the parameters of all distribution models; and an optimal hybrid model screening module, used to select the model with the optimal test statistic and no overfitting from the distribution models that have completed parameter solving based on KS hypothesis testing and Occam's razor principle, i.e., to give the number of branches of the optimal hybrid model and the statistical distribution type of each branch; the model parameter solving module is specifically used to: automatically solve the parameters of all distribution models using the expectation-maximization algorithm (EM); and the optimal hybrid model screening module is specifically used to: select the model whose p-value meets the requirements through KS hypothesis testing; and select the model with the smallest number of branches from the models whose p-value meets the requirements according to Occam's razor principle.

4. The estimation system for the optimal hybrid distribution parameters of vehicle-to-everything (V2X) communication opportunity intervals as described in claim 3, characterized in that, The optimal hybrid model selection module is specifically used to: define the model effectiveness index (MAC). p And the model simplicity index g: in, Representation Model The effectiveness index value; sce represents the scenario under different conditions, and there are a total of S scenarios; This indicates that in the scenario SCE, the model In optimal parameters The maximum p-value that can be obtained is expressed as: The value of g represents the number of branches in the mixture model; the higher the number of branches, the more complex the model. An overall metric G is introduced as follows: in, Representation Model The overall index value; Representation Model The simplicity index value; thus, the optimization problem for solving the optimal mixed distribution model is as follows: in, This represents the set of all possible mixed distribution models; by solving the optimization problem, a model with the optimal test statistic and without overfitting is obtained.