A method for calculating the association of vehicle codes and the association of face codes in electricity

By scanning CDR data and combining the acquisition time and location of image data, a weighted association rule mining algorithm and feature engineering were used to solve the problem of spatiotemporal inconsistency in the association between vehicle code and face code, thereby improving the accuracy and efficiency of the association.

CN116842432BActive Publication Date: 2026-02-06CHENGDU HELIO INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310351107.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2026-02-06
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

When existing technologies associate vehicle codes and facial recognition codes, the spatiotemporal inconsistency between image data and CDR data results in a large amount of accompanying IMSI code data and a large number of invalid associations, making it difficult to identify valid vehicle code and facial recognition code associations.

Method used

Using scanned CDR data, traditional association rule mining algorithms are used to calculate the associations between IMSI codes. Combining the acquisition time and location of image data, a weighted association rule mining algorithm and feature engineering are used to calculate the confidence and boost of candidate associations. A pre-trained binary classification model is used to calculate the association probability to reduce the impact of invalid associations.

Benefits of technology

It improves the accuracy of vehicle code and face code association, reduces the number of invalid associations, reduces the impact on activity areas and permanent personnel, avoids the impact of trajectory completion errors and base station trajectory deviations, and achieves efficient association calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The application discloses a kind of methods for vehicle code association and face code association in electric image calculation in the technical field of electric image calculation, including the CDR record of close acquisition time, adjacent base station spatial position as IMSI code accompanying record;Calculate the code code association between IMSI code;Image data is formed with image acquisition time close, with image acquisition rod body position adjacent base station connected each CDR record vehicle code, face code accompanying record;According to the number of CDR records accompanied by this image data, the number of code code association in CDR accompanying record determines the weight of accompanying record;Using the association rule mining algorithm with weight to mine out the candidate association of vehicle code and face code, record the confidence, promotion degree of each candidate association;Through feature engineering, the spatiotemporal range characteristics of the candidate association are calculated;According to the pre-trained binary classification model, the spatiotemporal range characteristics of the candidate association, the original confidence, promotion degree data of the candidate association, the association probability of each candidate association is calculated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric image calculation, and in particular to a method for vehicle code association and face code association in electric image calculation. BACKGROUND

[0002] Electric image calculation is a technology that fuses data from intelligent terminal sensing sources such as mobile phone IMSI code and IMEI code with image data such as human face pictures and license plate pictures through big data and artificial intelligence analysis, and then establishes the association relationship and behavior mode of related data source objects.

[0003] Among them, CDR (Call Detail Record) data is one of the main data sources in electric image calculation due to its low sensitivity. CDR is a record generated by mobile devices in user calls, SMS, Internet access and other behaviors, which records mobile phone IMSI code, collection time, user behavior and the geographic location of the connected base station and its cell.

[0004] Using CDR data to calculate vehicle code association and face code association has the following technical problems:

[0005] 1. The location area of the mobile phone can only be obtained from the mobile phone CDR data, and the accurate location cannot be obtained

[0006] The current location data of the CDR of the telecommunications operation data is the location of the base station connected by the mobile phone and the cell. The CDR data can only determine a cell range (Cell). The location of the mobile phone corresponding to the signaling data in the overlapping coverage range of multiple mobile phone base stations is more uncertain. At this time, the logical cell location of the CDR is not the same as the actual physical location of the mobile phone.

[0007] 2. The time and space of the face, vehicle and mobile phone code data are not synchronized

[0008] The locations of the vehicle and the face image are not consistent, and the time points of collection are also different. The relationship between the location data in the mobile phone CDR and the above two is range overlap, and the time and space are not synchronized, so the overlap in time is only in the range.

[0009] 3. Interference data brings a large number of invalid associations

[0010] Since the CDR data is a regional positioning, the CDR data of the same mobile phone in a time range overlaps with the CDR data of a large number of other mobile phones in the region or trajectory (such as: the same cell residents). At this time, if a mobile phone exists in association with a license plate and a face in this time range, then this batch of mobile phones all exist in association with the license plate and the face.

[0011] 4. The data collected from personnel and vehicles is not continuous at the collection location, and it cannot be guaranteed that there will be image data for personnel and vehicles that appear near the camera.

[0012] There are currently two main technical solutions for linking vehicle codes and facial recognition codes using CDR data and image data:

[0013] 1. Collision probability-based methods

[0014] This method extends upon traditional vehicle code association and face code association methods. Taking vehicle code association as an example, traditional vehicle code association calculates the number of times the IMSI code collected by the fake base station on the same pole appears alongside the vehicle within a time period t before and after the time the vehicle was photographed, based on the time the camera captured the vehicle and the pole's position. If the number of times the vehicle appears alongside one or more IMSI codes exceeds a specified threshold, the vehicle code association is considered successful. The time period t is used because the actual time the fake base station connects to the mobile device's IMSI code is inconsistent with the time of the camera capture, resulting in a certain time difference.

[0015] When this method is extended to CDR data for vehicle code or face code collision, considering that CDR data and camera data are not synchronized in time and space, the CDR data selected for collision with vehicles and faces needs to be within a certain time and space range. The space range needs to cover multiple adjacent base stations of the camera pole. At this time, the number of IMSI codes participating in the collision will be far greater than that of traditional vehicle code and face code association.

[0016] 2. Trajectory Similarity-Based Methods

[0017] This method aggregates vehicle, face, and IMSI code records, chronologically grouping records of the same vehicle, face, and IMSI code to form trajectory data for each vehicle, face, and IMSI code over a period of time. Then, it calculates the similarity between each pair of vehicle and IMSI code trajectories, and between face and IMSI code trajectories. Vehicle-to-code and face-to-code associations are considered successful if the trajectory similarity exceeds a certain threshold. Higher trajectory similarity increases the probability of association between vehicle-to-code and face-to-code. The method for calculating trajectory similarity varies depending on the chosen similarity calculation method.

[0018] The shortcomings of the existing technical solutions are described as follows:

[0019] 1. Deficiencies of collision probability-based methods

[0020] First, since the number of IMSI codes involved in each collision is very large, after multiple collisions with the IMSI codes that accompany the same vehicle or face, a large number of IMSI codes will still be associated with the vehicle and face. Increasing the threshold will cause the number of associated IMSI codes to drop sharply, and effective associations may be eliminated at the same time, resulting in collision failure.

[0021] Secondly, when calculating the association between vehicle code and face code, the proportion of CDR data that can effectively distinguish the location differences of the IMSI codes of the main activity areas and the trajectories of the people is low, resulting in insufficient discrimination during collision calculation and failure to identify effective associations.

[0022] Furthermore, the collision probability method using a uniform threshold cannot handle scenarios where vehicle codes and face codes of permanent and non-permanent personnel have different distribution probabilities. This is because the number of vehicle codes and face codes associated with non-permanent personnel is much lower than that associated with permanent personnel. If the threshold is too high, collisions with non-permanent personnel will fail, and if the threshold is too low, a large number of invalid associations associated with permanent personnel will occur.

[0023] Finally, the collision probability-based method only considers the contribution of the number of collisions to the association calculation, without considering the contribution of the temporal and spatial range of collision occurrences. Theoretical analysis reveals that for vehicle code and face code association, the larger the temporal and spatial range of collision occurrences, the greater the probability of a strong association between the vehicle code pair and the associated code pair.

[0024] 2. Limitations of trajectory similarity-based methods

[0025] First, due to factors such as camera acquisition rate, shooting angle, line of sight obstruction, and recognition errors, the image acquisition system cannot guarantee that vehicles and faces will be captured and correctly recognized every time the camera appears. Therefore, the trajectory of vehicles and faces is not a complete trajectory data and needs to be completed. However, the deviation between the completed trajectory and the actual trajectory will greatly affect the calculation results of trajectory similarity.

[0026] Secondly, since the location data in CDR data is base station data, and the trajectory data of the IMSI code is the trajectory of the phone connecting to the base station, not the actual trajectory of the phone, this trajectory does not overlap with the trajectory of the owner's vehicle or face, and the similarity is much lower than that of conventional trajectory similarity. If we consider the ping-pong effect of switching base stations back and forth when the CDR data overlaps with the base station coverage, the deviation between the base station trajectory formed by the IMSI code and the actual trajectory of the phone will further affect the trajectory similarity calculation results. Summary of the Invention

[0027] The technical problem this invention aims to solve is that, during correlation operations, due to the spatiotemporal inconsistency between image data and CDR data acquisition, the amount of IMSI code data accompanying vehicles and faces is large, and there are a large number of IMSI codes with similar activity areas and trajectories. The goal is to remove invalid accompanying data, identify the strong correlation between valid vehicle codes and face codes, and ultimately achieve the direct use of CDR data and vehicle and face images to complete the correlation operation of vehicle codes and face codes.

[0028] To solve the above technical problems, the application provides a method for calculating code association of vehicles and code association of faces, which adopts the following technical scheme:

[0029] Step one: scanning CDR data, and taking CDR records with close collection time and adjacent base station space position as IMSI code accompanying records;

[0030] Step two: calculating code association between IMSI codes by using a traditional association rule mining algorithm;

[0031] Step three: according to the collection time and pole position of vehicle and face image data, forming vehicle code and face code accompanying records of each CDR record connected by the base station with close collection time and adjacent pole position;

[0032] Step four: according to the total number of IMSI codes accompanying the image records, denoted as C1, and the number of code association among the IMSI codes, denoted as C2, calculating the weight w of the accompanying records, w is negatively related to C1 and C2;

[0033] Step five: using a weighted association rule mining algorithm to mine candidate associations of vehicle code and face code, and recording the confidence and lift of each candidate association;

[0034] Step six: for each candidate association, calculating the spatiotemporal range feature of the candidate association through feature engineering;

[0035] Step seven: according to the pre-trained binary classification model, the features calculated in step six, the candidate association confidence and lift data recorded in step five, calculating the association probability of each candidate association; if the association probability is higher than a specified threshold, the association is effective, otherwise, the association is invalid.

[0036] Optionally, in step one, the IMSI code accompanying record calculation process adopts a time slicing method to determine whether the CDR records are close in time, and the length of the time slice tsl is determined according to the statistical features of the CDR record stay duration.

[0037] Optionally, the IMSI code accompanying record calculation process adopts a space slicing method to determine whether the CDR records are adjacent in position, and the size of the space slice is determined according to the statistical features of the slice time t sl and the estimated travel speed v of the CDR corresponding mobile phone.

[0038] Optionally, in step one, the CDR data and image data are preprocessed, and the specific steps include:

[0039] I) merging continuous CDR data of the same mobile phone located in the same base station in the same mobile phone CDR data;

[0040] II) Merge ping-pong handover data from the same mobile CDR;

[0041] III) Identify stationary CDR records. A stationary CDR refers to a record where the same IMSI code remains continuously at the same base station for a period exceeding a preset threshold t. max This indicates that the user's activity range for an extended period is limited to the coverage area of ​​a single base station.

[0042] In step four, when calculating the accompanying record weights for vehicle and face codes, the weight w of the accompanying record for the static CDR data is... s proportionally r s Decrease, where: r s ∈[0,1],

[0043] w s =w*r s

[0044] IV) Remove drift data of the same subject. Drift data refers to the significant deviation in position of the same subject at different times. The subject can be IMSI code, vehicle, or human face.

[0045] V) Identify the mobile CDR records of resident personnel. Resident personnel are those who periodically appear repeatedly in the same time period and the same area within a certain time range.

[0046] After reducing the weight of static personnel in step III), the weight of the accompanying records of the mobile CDR data of resident personnel is w. r Further reduce proportionally, where: r r ∈[0,1),

[0047] w r =w s *r r .

[0048] Optionally, let the average shooting interval of the current camera be t. p The method for calculating the associated weighting ratio of the static CDR is as follows:

[0049] r s =t p / t max If t p <t max Otherwise r s =1;

[0050] Permanent staff weighting ratio r r The selection of the number of resident IMSI codes in the current region C r The negative correlation and the calculation method for the weighting ratio of permanent residents are as follows: r r =1 / ln(C r +1).

[0051] Optionally, in step three, the judgment method of close time collection is to determine the time interval according to the image collection time and the statistical characteristics of the CDR record stay duration; and the judgment method of the base station adjacent to the rod body position is to determine according to the image collection rod body position, the time interval and the vehicle speed or walking speed at the collection time.

[0052] Optionally, in step four, the weight calculation method of the accompanying record is: w = 1 / ln((C1+1)*(C2+1)).

[0053] Optionally, in steps one to six, the operation candidate association and its characteristics are supported in units of days, then the data of the recent period of time are weighted and summarized to obtain the candidate association and its characteristic data of the current period, and then the classification model calculation of step seven is performed.

[0054] Optionally, the space-time range characteristics include: total number of accompanying, number of accompanying rod bodies, maximum distance between accompanying rod bodies, maximum number of accompanying single rod bodies, average number of accompanying rod bodies, extreme difference of rod body accompanying number, standard deviation of rod body accompanying number, total number of days of accompanying occurrence, daily maximum accompanying number, daily average accompanying number, and daily accompanying number standard deviation.

[0055] Optionally, the process of the pre-trained binary classification model is as follows:

[0056] Referring to the above steps one to six, the candidate association and the characteristics of each candidate association of the current time in the region are generated;

[0057] According to the characteristics of the candidate association, the candidate association confidence generated in step five, the promotion degree, and the actual association result of the candidate association, a binary classification model of the region is trained using a logistic regression model or a fully connected neural network.

[0058] In summary, the present application includes at least one of the following beneficial effects:

[0059] 1. The method only returns candidate associations through collision probability calculation, and then performs secondary operation through feature engineering and a pre-trained classification model, so as to reduce the collision threshold and avoid collision failure caused by too high threshold;

[0060] 2. The candidate association is calculated using a weighted association rule mining algorithm, so as to reduce the influence of active region and active trajectory similar personnel, static personnel and permanent personnel on association calculation;

[0061] 3. The weight of the accompanying record of the permanent personnel and the static personnel is reduced, a uniform threshold can be set for the non-permanent personnel, the association calculation is uniformly performed, and the workload of operation is reduced;

[0062] 4. In addition to the number of occurrences, the spatiotemporal range, concentration, and dispersion of the spatiotemporal range can all be included in the correlation calculation, which greatly improves the accuracy of the correlation calculation.

[0063] 5. It does not require the formation of complete vehicle and face trajectories, thus avoiding the impact of missed data collection, recognition errors, and trajectory completion errors on association calculations; it also avoids the failure of similarity calculations caused by the deviation between base station trajectories and mobile phone trajectories. Detailed Implementation

[0064] The present invention will be further described in detail below.

[0065] Example 1

[0066] This invention discloses a method for associating vehicle codes and face codes in image processing, comprising the following steps:

[0067] Step 1: Scan the CDR data and use CDR records with similar acquisition times and adjacent base station spatial locations as IMSI code accompanying records;

[0068] Tid Accompanying IMS code list Accompanying start time Accompanying end time Accompanying base station 1 I1, I2, ... TB1 TE1 S1_1,S2_1… 2 I1, I3, ... … … … 3 … … 4 I2, I3,… … … …

[0069] Step 2: Use traditional association rule mining algorithms such as Apriori and FPGrowth to calculate the code-code associations between IMSI codes;

[0070] Step 3: Based on the acquisition time and pole position of the vehicle and face image data, form a vehicle code and face code accompanying record for each CDR record connected to the base station that is close to the acquisition time of this image data and adjacent to the pole position. Each accompanying record includes a vehicle (or face), an IMSI code and the weight of this record; the original record ID is used for subsequent reverse lookup of the original record to facilitate subsequent calculations.

[0071]

[0072] Step 4: Based on the total number of IMSI codes accompanying the same original image record, denoted as C1, and the number of associated codes among the accompanying IMSI codes, denoted as C2, calculate the weight w of the accompanying record; w is negatively correlated with C1 and C2; to avoid the weight decreasing too quickly due to excessively large C1 and C2 values, an optional weight calculation method is as follows:

[0073] w = 1 / ln((C1+1)*(C2+1))

[0074] The addition of 1 to both C1 and C2 is for smoothing, to prevent the weight formula from becoming invalid when there is only one companion or no code association.

[0075] Step five: using Weighted FpGrowth and other weighted association rule mining algorithms to mine candidate associations of vehicle code and face code, and record the confidence and lift of each candidate association;

[0076] Step six: for each candidate association, calculate the spatio-temporal range features of the candidate association through feature engineering, including but not limited to: total number of accompanying, number of accompanying rods, maximum distance between accompanying rods, maximum number of accompanying rods, average number of accompanying rods, range of accompanying rods, standard deviation of accompanying rods, total number of days of accompanying, daily maximum number of accompanying, daily average number of accompanying, daily standard deviation of accompanying;

[0077] Step seven: according to the pre-trained binary classification model, the features calculated in step six, the candidate association confidence and lift data recorded in step five, calculate the association probability of each candidate association; higher than the specified threshold is an effective association, otherwise it is an invalid association, in general the threshold can be 0.5, or it can be set to other thresholds according to business needs, such as increasing the precision of the association or reducing the threshold to improve the recall rate of the association.

[0078] The pre-trained model required in step seven needs to rely on the existing vehicle code association and face code association result data in the region and the same period CDR data, vehicle and face camera data for model training, the process is as follows:

[0079] Refer to steps 1-6 above to generate candidate associations and features of each candidate association in the region at this time;

[0080] According to the features of the candidate associations calculated in step 6, the confidence and lift of the candidate associations generated in step 5, and the actual association results (yes or no) of the candidate associations, use a logistic regression model (or a fully connected neural network) to train a binary classification model for the region.

[0081] As another preferred embodiment of the present embodiment:

[0082] In the process of calculating IMSI code accompanying records, time slicing can be used to determine whether the CDR records are time close, and the length of the time slicing is determined according to the statistical characteristics of the CDR record stay time, the specific steps are as follows:

[0083] 1) Sort the CDR records in ascending order according to IMSI code and collection time;

[0084] 2) Calculate the current stay time of the IMSI code in the CDR record, record the collection time of the CDR record as T, and the collection time of the next record of the same mobile phone as T n , the current stay time t st = T n-T, the total dwell time interval of this CDR record can be denoted as [T, T+t]. st ];

[0085] 3) Calculate the slice time t based on the statistical characteristics of the dwell time. sl For each day, time is divided into slices t. sl Slicing is performed, and the slice size can be selected from the median m of the dwell time, the mean μ, the upper quartile q3, the upper limit of anomaly detection using the quartile method: q3 + 1.5*(q3 - q1), etc. If it is necessary to construct the spatiotemporal association relationship of the IMSI code to the maximum extent, t can be selected. sl = q3 + 1.5 * (q3 - q1). To avoid omissions in the calculation of CDR data appearing at the edge of the time slice, each time slice is slid according to t. sl / 2 slides the window, at which point a time slice of length t will appear between adjacent time slices. sl The overlap time is 2 / 2, and the specific slices are as follows:

[0086] [-t sl / 2,t sl / 2],[0,t sl ],…,[k*t sl / 2,k*t sl / 2+t sl ],……, where: k≥-1

[0087] negative number -t sl / 2 represents t before 00:00 AM. sl The time of / 2 indicates that the record at the beginning of each day needs to be calculated together with the record from the previous day near midnight.

[0088] 4) Record the dwell interval [T, T+t] in the CDR. st ] and each time slice after slicing [k*t sl / 2,k*t sl / 2+t sl The two are compared, and if they intersect, it means that the CDR record falls within the time slice. The IMSI code corresponding to the CDR record in the same time slice forms an accompaniment. Generally, a CDR record will appear in multiple time slices.

[0089] During the calculation of IMSI code-accompanying records, spatial slicing is used to determine whether CDR records are adjacent in terms of spatial location of base stations. The size of the spatial slicing is based on the slice time t. sl The statistical characteristics of the estimated travel speed v of the mobile phone corresponding to CDR are determined, and the specific steps are as follows:

[0090] 1) Sort the CDR records in ascending order by IMSI code and acquisition time;

[0091] 2) Estimate the travel speed v of the corresponding mobile phone in the CDR record, record the collection time of the CDR record as T, record the base station position as P, and record the collection time of the nearest CDR record of the same mobile phone at different base station positions as T n2 , record the base station position as P n , and record the base station stay time t st2 = T n2 -T, calculate the distance s n between base stations P and P n , the travel speed v of the mobile phone = s n / t st2 ;

[0092] 3) According to the statistical characteristics of the travel speed v s and the slicing time t sl , calculate the distance s sl= v s* t sl , v s may select the median m, mean μ, upper 4 quantile q3, 4 quantile method upper limit: q3+1.5*(q3-q1) of the travel speed of the mobile phone. Use s sl as the window size to slice the corresponding spatial range of the CDR record set, and set the moving window of the slice to s sl / 2 to ensure the integrity of the accompanying records of the edge records of the slice. The specific operation of the spatial range slicing includes the following steps:

[0093] a) Calculate the minimum value lon min , maximum value lon max , minimum value lat min and maximum value lat max of the longitude of the base station position in the CDR record set;

[0094] b) Slice calculation in the order of longitude first and latitude second, latitude lat min , starting from the minimum value of the base station longitude lon min , the diagonal vertex coordinates of the spatial area at this time are [(lat min lon min ), (lat min +s sl lon min +s sl )], move in the longitude direction by s sl / 2 each time, and the diagonal vertex coordinates of the spatial range after moving are [(lat min lon min +s sl / 2), (latmin + sl lon min + sl / 2 sl ], until the maximum longitude lon max is crossed;

[0095] c) latitude is moved by displacement s sl / 2, longitude back to the minimum value lon min again, the diagonal vertex coordinates of the spatial region at this time are [(lat min + sl / 2 lon min ), (lat min + sl + sl / 2 lon min + sl )], each time moving in the longitude direction by displacement s sl / 2, the diagonal vertex coordinates of the spatial region after moving are [(lat min + sl / 2 lon min + sl / 2), (lat minn + sl + sl / 2 lon min + sl / 2 + s sl ], until the maximum longitude lon max is crossed at the current latitude;

[0096] d) repeat c until the latitude crosses the maximum latitude lat max ;

[0097] 4) compare the base station position P in the CDR data with the space slice S r,c : [(lat min + r * s sl / 2 lon min + c * s sl / 2), (lat min + r * s sl / 2 + s sl lon min + c * s sl / 2 + s sl ] are compared, r, c represent row and column numbers, if P is covered by the space slice S r,c , it indicates that P falls in the space slice S r,cWithin the same spatial slice, the IMSI code corresponding to a CDR record forms an accompaniment; generally, a CDR record will appear in multiple spatial slices.

[0098] As another preferred embodiment of this embodiment:

[0099] CDR recordings with acquisition and image recording times close to each other (T2) and adjacent camera pole positions (P2) are formed simultaneously.

[0100] The collection time can be selected similarly based on the statistical characteristics of the dwell time in the CDR records, such as: the median m, mean μ, upper quartile q3, and the upper limit of anomaly detection using the quartile method: q3 + 1.5*(q3 - q1), etc. To construct the spatiotemporal association relationship to the maximum extent, t can be selected. sl = q3 + 1.5*(q3 - q1), the time interval with similar collection times can be determined as [T2 - t sl / 2,T2+t sl / 2], this time interval is compared with the total dwell interval [T,T+t] of each CDR record. st The two are compared, and if they intersect, they are considered to be close in time.

[0101] The spatial proximity of base stations is not limited to the geographical proximity of the base station and pole location P2 in the CDR record. Records where the distance between the CDR base station and pole P2 is less than a certain threshold s2 can be selected. In this case, multiple base stations may be involved. The calculation of s2 is related to the vehicle speed or walking speed at the time of data collection. Assuming the vehicle speed or walking speed at the time of data collection is v2, then s... 2= v 2* t sl Taking vehicle images as an example, the estimation method for v2 is as follows:

[0102] 1) Sort the vehicle video recordings in ascending order by license plate number and collection time;

[0103] 2) Estimate the vehicle's speed v2: Let the image recording acquisition time be T2, the camera pole position be P2, and the acquisition time of the next record for the same license plate be T. 2n The position of the rod is denoted as P. 2n The duration of the rod's stay is t st2 =T 2n -T2, calculate the relationship between the two rods P2 and P... 2n distance s p The mobile phone's speed is v2 = s p / t st2 ;

[0104] As another preferred embodiment of this embodiment:

[0105] In step one, the CDR data and video data are preprocessed. Through preprocessing, the data volume can be greatly reduced, the quality of the CDR data can be improved, and the algorithm efficiency can be improved. The preprocessing steps include:

[0106] I) merging the continuous CDR data of the same mobile phone CDR data located in the same base station, and the specific steps include:

[0107] a) arranging the CDR records in ascending order according to the IMSI code and the collection time;

[0108] b) calculating the distance s_prev between the CDR record and the previous CDR record of the mobile phone, the stay duration ts_prev, the distance s_next between the CDR record and the next CDR record of the mobile phone, and the stay duration ts_next;

[0109] c) s_prev=0 indicates that the record and the previous record have the same position, and is deleted;

[0110] d) performing steps a) and b) again to calculate s_prev, ts_prev, s_next, and ts_next of the merged record;

[0111] II) merging the ping-pong switching data of the same mobile phone CDR data. The CDR ping-pong switching data refers to the position data of the CDR record jumping back and forth between two cells when the mobile phone is in the overlapping area of the base stations, and the base station position is displayed as A->B->A.

[0112] a) arranging the CDR records processed in step I in ascending order according to the IMSI code and the collection time;

[0113] b) calculating the base station position sta_prev of the previous record of the IMSI code in the CDR record, and calculating the base station position sta_next of the next record of the IMSI code;

[0114] c) comparing the previous and next stay positions sta_prev and sta_next of the CDR record. If the previous and next positions are the same, the stay times ts_prev and ts_next are compared. If any of the stay times is less than a preset threshold t sw , the data is considered as ping-pong switching data, otherwise, it is considered as actual stay in both base stations;

[0115] d) according to a preset rule, the position of the ping-pong switching data is uniformly switched to the target position data. For example, the rule can select the record with the smallest base station table sequence number in the current ping-pong switching;

[0116] e) merging the continuous records of the same IMSI code again in the manner of step I;

[0117] III) Identify the CDR record of the static state of the mobile phone, the static state CDR refers to the continuous residence time of the same IMSI code in the same base station exceeding the preset threshold t max (For example: 30 minutes), indicating that the machine owner's activity range is limited to the coverage range of a single base station for a long time, which can basically be regarded as the machine owner being in a static state. In the calculation of the vehicle code and face code, the static state CDR will have a large number of accompanying records with the image records of adjacent rods in the static time. Such accompanying records should have a contribution degree much lower than that of the conventional CDR accompanying records when calculating the vehicle code and face code correlation.

[0118] In step four, when calculating the accompanying record weight of the vehicle code and face code, the weight of the accompanying record of the static state CDR data needs to be reduced by a proportion r s , where: r s ∈[0,1].

[0119] w s =w*r s

[0120] If the accompanying of the static state is not considered, r s =0; if the IMSI code accompanying of the static state needs to be considered, set the average camera interval time of the current camera as t p , then:

[0121] r s =t p / t s If t p <t s , otherwise r s =1

[0122] IV) Remove the drift data of the same subject, which refers to the obvious deviation of the position of the same subject at different times, which may be caused by system abnormalities or vehicle license plate and face recognition errors. The subject can be an IMSI code, a vehicle, or a face. The specific steps are as follows:

[0123] a) Sort the various data according to the same subject and in ascending order of time;

[0124] b) Calculate the distance s_prev of the current record from the previous record of the subject, the time ts_prev, and v_prev=s_prev / t_prev;

[0125] c) Calculate the distance s_next of the current record from the next record of the subject, the time ts_next, and v_next=s_next / t_next

[0126] d) If v_prev and v_next are both greater than the preset threshold v max , it is considered as drift record and directly deleted;

[0127] If the same subject appears in different locations at the same time, the drift data can be determined by referring to the previous and subsequent positions of the same subject, and then removed.

[0128] V) Identify the CDR records of the mobile phone of the resident personnel, who are personnel who appear repeatedly in the same time period and the same area within a certain time range. The specific steps include:

[0129] a) Calculate the total time and proportion of each IMSI code in each base station position in the CDR records within a certain period of time (e.g., 1 month);

[0130] b) If the proportion is higher than the preset threshold r max or the top n cell positions in terms of proportion are considered as the resident area of the mobile phone;

[0131] c) Identify the CDR data of the mobile phone resident area;

[0132] d) After reducing the weight of the static CDR in step III), the resident personnel will form a large number of accompanying records with the image data of the surrounding cameras, and the weight of the accompanying records of the resident personnel CDR data needs to be further reduced by a proportion r r , where r r ∈[0,1).

[0133] w r =w s *r r ,

[0134] If the accompanying of the resident personnel is not considered, r r =0; if the accompanying of the resident personnel is considered, the selection of r r should be negatively related to the number of resident IMSI codes C r in the current area, and a selectable calculation scheme is as follows:

[0135] r r =1 / ln(C r +1).

[0136] As another preferred embodiment of the present embodiment:

[0137] In steps one to six, daily operation of candidate association and its characteristics is supported, and then the data of the last period of time (e.g., 30 days) is weighted and summarized to obtain the candidate association and its characteristic data of the current period, and then the classification model calculation of step seven is performed; the daily calculation result record of candidate association is as follows:

[0138]

[0139]

[0140] The invention point of the method is:

[0141] 1. The unsupervised learning association rule mining algorithm is combined with the supervised learning classification algorithm, the vehicle code and the face code are associated in a wider time and space range, the accompanying number and the accompanying time and space range are included in the association probability operation, and the problem of collision failure caused by too many accompanying records is effectively avoided

[0142] 2. The weighted association rule mining algorithm is used to calculate the candidate association, the influence of the active area, the active track similar personnel, the static personnel and the resident personnel on the association calculation is reduced, and the main calculation factors of the weight include:

[0143] a) When the image is collected, the flow of people around the collection point is used to represent the total number of accompanying IMSI codes;

[0144] b) When the image is collected, the total number of mutually associated IMSI codes in the accompanying IMSI code;

[0145] c) When the image is collected, whether the accompanying IMSI code is in a static state;

[0146] d) When the image is collected, whether the accompanying IMSI code belongs to a resident;

[0147] 3. The characteristics of the candidate association calculated through the feature engineering include the number of times, the time range, the space range, the concentration and the dispersion of the time and space range of the accompanying occurrence of the candidate association, and the pre-trained binary classification model is used to calculate the association probability, so that the accuracy of the association calculation can be improved.

[0148] The above are preferred embodiments of the application, and are not intended to limit the protection scope of the application, therefore: any equivalent changes made on the basis of the structure, shape and principle of the application should be covered within the protection scope of the application.

Claims

1. A method for associating vehicle codes and face codes in image processing, characterized in that: Includes the following steps: Step 1: Scan the CDR data and use CDR records with similar acquisition times and adjacent base station spatial locations as IMSI code accompanying records; Step 2: Treat all IMSI code-related records generated in Step 1 as a transaction database, and use an association rule mining algorithm to calculate the code-code associations between IMSI codes; Step 3: Based on the acquisition time and pole position of the vehicle and face image data, form a vehicle code and face code accompanying record for each CDR record connected to the base station that is close to the acquisition time of this image data and adjacent to the pole position; Step 4: Based on the total number of IMSI codes accompanying the image record, denoted as C1, and the number of code associations among the accompanying IMSI codes, denoted as C2, calculate the weight w of the accompanying record, where w is negatively correlated with C1 and C2. In step four, the weight of the accompanying record is calculated as follows: w = 1 / ln((C1+1)*(C2+1)); Step 5: Use a weighted association rule mining algorithm to mine candidate associations for vehicle QR codes and face codes, and record the confidence and lift of each candidate association; Step Six: For each candidate association, through feature engineering, calculate the spatiotemporal range features of the candidate association from the accompanying records of vehicle code and face code formed in Step Three; Step 7: Based on the pre-trained binary classification model, the features calculated in Step 6, and the candidate association confidence and lift data recorded in Step 5, calculate the association probability of each candidate association; if the association probability is higher than a specified threshold, it is a valid association; otherwise, it is an invalid association.

2. The method for vehicle code association and face code association in image processing according to claim 1, characterized in that: In step one, the IMSI code-accompanying record calculation process uses time-segmentation to determine whether CDR records are time-close. The time segment t... sl The length is determined based on the statistical characteristics of the dwell time recorded in the CDR.

3. The method for associating vehicle codes and face codes in image processing according to claim 2, characterized in that: In step one, the IMSI code-accompanying record calculation process uses spatial slicing to determine whether CDR records are adjacent. The size of the spatial slicing is based on the slice time t. sl And the statistical characteristics of the estimated travel speed v of the mobile phone corresponding to CDR are determined.

4. The method for vehicle code association and face code association in image processing according to claim 1, characterized in that: In step one, the CDR data and image data are preprocessed, and the specific steps include: I) Merge consecutive CDR data located at the same base station within the same mobile phone CDR data; II) Merge ping-pong handover data from the same mobile CDR; III) Identify stationary CDR records. A stationary CDR refers to a record where the same IMSI code remains continuously at the same base station for a period exceeding a preset threshold t. max This indicates that the user's activity range for an extended period is limited to the coverage area of ​​a single base station. In step four, when calculating the accompanying record weights for vehicle and face codes, the weight w of the accompanying record for the static CDR data is... s proportionally r s Decrease, where: r s ∈[0,1], w s =w*r s IV) Remove drift data of the same subject. Drift data refers to the significant deviation in position of the same subject at different times. The subject can be IMSI code, vehicle, or human face. V) Identify the mobile CDR records of resident personnel. Resident personnel are those who periodically appear repeatedly in the same time period and the same area within a certain time range. After reducing the weight of static personnel in step III), the weight of the accompanying records of the mobile CDR data of resident personnel is w. r Further reduce proportionally, where: r r ∈[0,1), w r =w s *r r 。 5. The method for vehicle code association and face code association in image calculation according to claim 4, characterized in that: Let the average camera recording interval be t. p The method for calculating the weighting ratio of the static mobile phone CDR is as follows: r s =t p / t max If t p <t max Otherwise r s =1; Permanent staff weighting ratio r r The selection of the number of resident IMSI codes in the current region C r The negative correlation and the calculation method for the weighting ratio of permanent residents are as follows: r r =1 / ln(C r +1).

6. The method for vehicle code association and face code association in image processing according to claim 1, characterized in that: In step three, the method for determining if the acquisition time is close is to determine the time interval based on the statistical characteristics of the image acquisition time and the duration of CDR recording; the method for determining if a base station is adjacent to the pole position is to determine the position of the image acquisition pole, the time interval, and the vehicle speed or walking speed at the acquisition time.

7. The method for vehicle code association and face code association in image processing according to claim 1, characterized in that: In steps one through six, it is possible to calculate candidate associations and their features on a daily basis, then weight and summarize the data from the most recent period to obtain the candidate associations and their feature data for the current period, and then perform the classification model calculation in step seven.

8. The method for vehicle code association and face code association in image processing according to claim 1, characterized in that: The spatiotemporal range characteristics include: total number of accompanying events, number of accompanying poles, maximum distance between accompanying poles, maximum number of accompanying events per pole, average number of accompanying events per pole, range of accompanying events per pole, standard deviation of accompanying events per pole, total number of days with accompanying events, maximum number of accompanying events per day, average number of accompanying events per day, and standard deviation of accompanying events per day.

9. The method for vehicle code association and face code association in image processing according to claim 1, characterized in that: The pre-training process for the binary classification model is as follows: Referring to steps one and six above, generate candidate associations for this region at this time and the characteristics of each candidate association; Based on the characteristics of the candidate associations, the confidence and lift of the candidate associations generated in step five, and the actual association results of the candidate associations, a binary classification model for this region is trained using a logistic regression model or a fully connected neural network.

Citation Information

Patent Citations

  • Method and device for associating characteristic data, storage medium and terminal

    CN109325429A

  • Time slice division method and system for adjoint analysis

    CN111242247A