Risk identification method and system based on video data and mobile phone signaling analysis

By performing video semantic mining and trajectory semantic mining on surveillance video data and mobile signaling data, a video trajectory aggregation vector is formed, which solves the problem of low reliability of risk identification in existing technologies and achieves more reliable and accurate risk identification.

CN121661571BActive Publication Date: 2026-04-17BAZHONG DATA GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BAZHONG DATA GROUP CO LTD
Filing Date
2026-02-04
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, mobile signaling data does not fully utilize potential semantic information in risk identification, resulting in relatively low reliability of risk identification.

Method used

By acquiring surveillance video data and mobile phone signaling data, video semantic mining and trajectory semantic mining are performed to form video semantic vectors and trajectory semantic vectors. These vectors are then aggregated to form video trajectory aggregate vectors, which are ultimately used for risk identification.

Benefits of technology

It improved the reliability and accuracy of risk identification, enriched the basis for risk identification, and enhanced the accuracy of risk identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661571B_ABST
    Figure CN121661571B_ABST
Patent Text Reader

Abstract

The application provides a risk identification method and system based on video data and mobile phone signaling analysis, and relates to the technical field of data analysis.In the application, first, monitoring video data and mobile phone signaling data formed in a target time interval in a target area are acquired, and motion trajectory data of each mobile phone user is determined based on the mobile phone signaling data;second, video semantic mining is performed on the monitoring video data to form a video semantic vector;then, trajectory semantic mining is performed on the motion trajectory data to form a trajectory semantic vector, wherein different trajectories have different mining methods in the trajectory semantic mining process;further, the video semantic vector and the trajectory semantic vector are aggregated to form a video trajectory aggregation vector;finally, risk identification is performed based on the video trajectory aggregation vector to obtain a risk identification result.Based on the above method, the problem of relatively low reliability of risk identification in the prior art can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, and more specifically, to a risk identification method and system based on video data and mobile phone signaling analysis. Background Technology

[0002] In traditional smart city construction, front-end sensing largely focuses on "seeing," with cameras and various sensors primarily handling image acquisition and signal uploading. This reliance on manual analysis is inefficient and prone to omissions. However, with the embedding of artificial intelligence algorithms into front-end devices, video surveillance can automatically perform crowd density analysis, target recognition, and abnormal behavior detection. New sensing methods such as mobile phone signaling and IoT terminals have evolved from simple counting to intelligent understanding of pedestrian flow and equipment status. The front end is no longer merely a passive "electronic eye" recording data, but is gradually becoming an "intelligent nerve ending" capable of automatic identification, preliminary judgment, and proactive alerts, providing more valuable structured information for back-end decision-making. However, the inventors' research has revealed that in existing technologies, mobile phone signaling is generally used as the basis for camera scheduling. For example, when a mobile phone signal indicates a high population density in an area, the camera is activated to collect video data, which is then analyzed to determine the level of risk. This approach fails to fully utilize the potential semantic information within the data, resulting in relatively low reliability of the identification. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a risk identification method and system based on video data and mobile signaling analysis, so as to improve the problem of relatively low reliability of risk identification in the prior art.

[0004] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:

[0005] A risk identification method based on video data and mobile signaling analysis includes:

[0006] Acquire surveillance video data and mobile phone signaling data generated in the target area within the target time interval, and determine the motion trajectory data of each mobile phone user based on the mobile phone signaling data;

[0007] Perform video semantic mining on the surveillance video data to form video semantic vectors;

[0008] The motion trajectory data is subjected to trajectory semantic mining to form trajectory semantic vectors. In the process of trajectory semantic mining, different trajectories with different complexities have different mining methods.

[0009] The video semantic vector and the trajectory semantic vector are aggregated to form a video trajectory aggregate vector;

[0010] Risk identification is performed based on the video trajectory aggregation vector to obtain risk identification results, wherein the risk identification results are used to reflect the degree of user behavior risk existing in the target area.

[0011] In some preferred embodiments, in the above-described risk identification method based on video data and mobile signaling analysis, the step of performing trajectory semantic mining on the motion trajectory data to form a trajectory semantic vector includes:

[0012] Each trajectory coordinate in each of the motion trajectory data is mapped to form a trajectory coordinate mapping parameter corresponding to each trajectory coordinate.

[0013] The trajectory coordinate mapping parameters corresponding to each trajectory coordinate in each of the motion trajectory data are combined to form an initial trajectory coordinate matrix. Then, based on a predetermined target matrix size, the initial trajectory coordinate matrix is ​​processed to form a target trajectory coordinate matrix, wherein the size of the initial trajectory coordinate matrix is ​​less than or equal to the target matrix size, and the size of the target trajectory coordinate matrix is ​​equal to the target matrix size.

[0014] If the trajectory complexity of the target trajectory coordinate matrix is ​​greater than the predetermined target complexity, then the target trajectory coordinate matrix is ​​subjected to multi-scale semantic mining and fusion to form a trajectory semantic vector, wherein the trajectory complexity is at least used to reflect the degree of overlap between different trajectories.

[0015] If the trajectory complexity of the target trajectory coordinate matrix is ​​not greater than the target complexity, then single-scale semantic mining is performed on the target trajectory coordinate matrix to form a trajectory semantic vector.

[0016] In some preferred embodiments, in the above-described risk identification method based on video data and mobile signaling analysis, the step of performing multi-scale semantic mining and fusion on the target trajectory coordinate matrix to form a trajectory semantic vector if the trajectory complexity of the target trajectory coordinate matrix is ​​greater than a predetermined target complexity includes:

[0017] The target trajectory coordinate matrix is ​​subjected to feature semantic extraction using multiple different semantic extraction sizes to form a trajectory extraction semantic vector corresponding to each semantic extraction size.

[0018] For each semantic extraction dimension belonging to the first dimension category, the trajectory extraction semantic vector is subjected to semantic compression to form a trajectory compressed semantic vector.

[0019] For each semantic extraction size belonging to the second size category, the trajectory extraction semantic vector is subjected to a semantic expansion operation to form a trajectory expansion semantic vector, wherein the size of the trajectory expansion semantic vector is equal to the size of the trajectory compression semantic vector.

[0020] Based on the semantic importance corresponding to each of the semantic extraction dimensions, each trajectory compressed semantic vector and each trajectory expanded semantic vector are semantically fused to form a trajectory semantic vector.

[0021] In some preferred embodiments, in the risk identification method based on video data and mobile signaling analysis described above, the step of semantically fusing each trajectory compressed semantic vector and each trajectory expanded semantic vector based on the semantic importance corresponding to each semantic extraction size to form a trajectory semantic vector includes:

[0022] Based on the semantic importance corresponding to each of the semantic extraction dimensions, each of the trajectory compressed semantic vectors and each of the trajectory expanded semantic vectors are semantically fused to form a composite dimension semantic vector.

[0023] The composite size semantic vector is deeply mined by using multiple dilated convolutional units with different receptive fields to form a trajectory depth semantic vector corresponding to each dilated convolutional unit.

[0024] Based on the semantic importance of each dilated convolutional unit, the trajectory depth semantic vector is semantically fused to form a trajectory semantic vector.

[0025] In some preferred embodiments, in the risk identification method based on video data and mobile signaling analysis described above, the step of semantically fusing each trajectory compressed semantic vector and each trajectory expanded semantic vector based on the semantic importance corresponding to each semantic extraction size to form a composite size semantic vector includes:

[0026] Linear mapping is performed on each of the trajectory compression semantic vectors and each of the trajectory expansion semantic vectors to form corresponding trajectory linear semantic vectors; and gating activation is performed on each of the trajectory linear semantic vectors to form a trajectory gating parameter distribution.

[0027] For each of the trajectory gating parameter distributions, the corresponding trajectory compression semantic vector or trajectory expansion semantic vector is adjusted based on the trajectory gating parameter distribution to form a trajectory adjustment semantic vector. Furthermore, a gating focus evaluation is performed on the trajectory gating parameter distribution to form a focus evaluation coefficient corresponding to the trajectory gating parameter distribution. The focus evaluation coefficient is used to reflect the uniformity of the distribution of each parameter in the trajectory gating parameter distribution.

[0028] For each semantic extraction size, the corresponding focus evaluation coefficient is adjusted based on the importance coefficient that is predetermined for the semantic importance corresponding to that semantic extraction size, to form the target weight parameter corresponding to that semantic extraction size;

[0029] Based on the target weight parameters corresponding to each semantic extraction size, each trajectory adjustment semantic vector is weighted and summed to form a composite size semantic vector.

[0030] In some preferred embodiments, in the above-described risk identification method based on video data and mobile signaling analysis, the step of semantically fusing each trajectory depth semantic vector based on the semantic importance corresponding to each dilated convolutional unit to form a trajectory semantic vector includes:

[0031] Each trajectory depth semantic vector is subjected to self-attention processing to form a corresponding trajectory self-attention semantic vector. Attention focus evaluation is performed based on the attention parameter distribution during the self-attention processing to obtain the focus evaluation coefficient corresponding to each trajectory self-attention semantic vector. The focus evaluation coefficient is used to reflect the uniformity of the distribution of each parameter in the attention parameter distribution.

[0032] For each of the dilated convolutional units, the corresponding focus evaluation coefficients are adjusted based on the importance coefficients that are predetermined for the semantic importance of the dilated convolutional unit, to form the target weight parameters corresponding to the dilated convolutional unit.

[0033] Based on the target weight parameters corresponding to each dilated convolutional unit, each trajectory self-attention semantic vector is weighted and summed to form a trajectory semantic vector.

[0034] In some preferred embodiments, in the above-described risk identification method based on video data and mobile signaling analysis, the step of aggregating the video semantic vector and the trajectory semantic vector to form a video trajectory aggregation vector includes:

[0035] Attention aggregation is performed on the video semantic vector and the trajectory semantic vector to form a video trajectory attention vector;

[0036] The video trajectory attention vector and the video semantic vector are subjected to multiple stages of attention aggregation to form multiple stages of video trajectory depth vectors, wherein the object of attention aggregation in a later stage includes the video trajectory depth vector of the previous stage.

[0037] Based on the video trajectory depth vector of the last stage, the video trajectory aggregation vector is determined.

[0038] In some preferred embodiments, in the above-described risk identification method based on video data and mobile signaling analysis, the step of performing multi-stage attention aggregation on the video trajectory attention vector and the video semantic vector to form a multi-stage video trajectory depth vector includes:

[0039] In the first stage of attention aggregation, the video trajectory attention vector and the video semantic vector are aggregated to form the video trajectory depth vector of the first stage.

[0040] In each subsequent attention aggregation stage, attention aggregation is performed on the video trajectory depth vector and the video semantic vector from the previous stage to form the video trajectory depth vector for the current stage.

[0041] In some preferred embodiments, in the above-described risk identification method based on video data and mobile signaling analysis, the step of performing video semantic mining on the surveillance video data to form a video semantic vector includes:

[0042] Based on the frame rate of the monitoring video data, a target parameter is determined, and based on the target parameter, the monitoring video data is processed by frame extraction to form monitoring video frame-extracted data, and the differential video frames between every two adjacent monitoring video frames in the monitoring video frame-extracted data are determined respectively, wherein the target parameter and the frame rate have a positive correlation.

[0043] Semantic mining is performed on each frame of the surveillance video data to form a semantic vector for each surveillance video frame. Semantic mining is also performed on each frame of the differential video frame to form a semantic vector for each differential video frame.

[0044] Semantic fusion is performed on the semantic vector of each of the monitored video frames and the semantic vector of each of the differential video frames to form a video semantic vector.

[0045] This invention also provides a risk identification system based on video data and mobile signaling analysis, including a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned risk identification method based on video data and mobile signaling analysis.

[0046] The risk identification method and system based on video data and mobile signaling analysis provided in this invention first acquires surveillance video data and mobile signaling data generated in a target area within a target time interval, and determines the motion trajectory data of each mobile phone user based on the mobile signaling data. Second, video semantic mining is performed on the surveillance video data to form a video semantic vector. Then, trajectory semantic mining is performed on the motion trajectory data to form a trajectory semantic vector, wherein different trajectory complexities employ different mining methods during the trajectory semantic mining process. Further, the video semantic vector and the trajectory semantic vector are aggregated to form a video trajectory aggregation vector. Finally, risk identification is performed based on the video trajectory aggregation vector to obtain the risk identification result. Based on the above method, since both video semantic mining and trajectory semantic mining are performed, the potential semantic information from both video and trajectory dimensions can be mined and aggregated to complete risk identification, making the semantic information on which risk identification is based richer, thereby improving the reliability of risk identification. Furthermore, since trajectory complexity is also considered during trajectory semantic mining, the accuracy of mining potential semantic information in the trajectory dimension can be improved to a certain extent, thereby improving the accuracy of the basis for risk identification, further improving the reliability of risk identification results, and thus improving the problem of relatively low reliability of risk identification in existing technologies.

[0047] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0048] Figure 1 This is a structural block diagram of a risk identification system based on video data and mobile signaling analysis provided in an embodiment of the present invention.

[0049] Figure 2 This is a flowchart illustrating the steps of the risk identification method based on video data and mobile signaling analysis provided in this embodiment of the invention.

[0050] Figure 3 This is a schematic diagram of video semantic mining provided in an embodiment of the present invention.

[0051] Figure 4 This is a schematic diagram of trajectory semantic mining provided in an embodiment of the present invention.

[0052] Figure 5 This is a schematic diagram illustrating multi-scale semantic mining and fusion provided in an embodiment of the present invention.

[0053] Figure 6 This is a first schematic diagram of semantic fusion provided for an embodiment of the present invention.

[0054] Figure 7 This is a second schematic diagram of semantic fusion provided for an embodiment of the present invention.

[0055] Figure 8 This is a third schematic diagram of semantic fusion provided for an embodiment of the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0057] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0058] like Figure 1 As shown, this embodiment of the invention provides a risk identification system based on video data and mobile phone signaling analysis. The risk identification system based on video data and mobile phone signaling analysis may include a memory and a processor.

[0059] In detail, the memory and the processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, they can be electrically connected via one or more communication buses or signal lines. The memory may store at least one software functional module (computer program) that exists in the form of software or firmware. The processor can be used to execute the executable computer program stored in the memory, thereby implementing the risk identification method based on video data and mobile phone signaling analysis provided in the embodiments of the present invention (as described below).

[0060] Optionally, the memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The processor may be a general-purpose processor, including a Central Processing Unit (CPU), Network Processor (NP), System on Chip (SoC), etc.; it may also be a Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0061] and, Figure 1 The structure shown is for illustrative purposes only. The risk identification system based on video data and mobile signaling analysis may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for exchanging information with other devices (such as video surveillance equipment).

[0062] In one alternative example, the risk identification system based on video data and mobile signaling analysis can be a server with data processing capabilities.

[0063] Combination Figure 2 This invention also provides a risk identification method based on video data and mobile signaling analysis, which can be applied to the aforementioned risk identification system based on video data and mobile signaling analysis. The method steps defined in the process related to the risk identification method based on video data and mobile signaling analysis can be implemented by the risk identification system based on video data and mobile signaling analysis (hereinafter referred to as the risk identification system).

[0064] The following will be about Figure 2 The specific process shown will be explained in detail.

[0065] Step S110: Obtain surveillance video data and mobile phone signaling data generated in the target area within the target time interval, and determine the motion trajectory data of each mobile phone user based on the mobile phone signaling data.

[0066] In this embodiment of the invention, the risk identification system can acquire surveillance video data and mobile phone signaling data generated in the target area within a target time interval, and determine the movement trajectory data of each mobile phone user based on the mobile phone signaling data. That is, the target area can be monitored using one or more video surveillance devices to collect image information and form surveillance video data. Furthermore, the mobile phone signaling data can be location-related data, such as GPS positioning data and base station signaling (Cellular Network Signaling), enabling the determination of the movement trajectory data of the corresponding mobile phone user based on the mobile phone signaling data. It should also be noted that the determined movement trajectory data can refer to a trajectory located within the target area.

[0067] Step S120: Perform video semantic mining on the monitoring video data to form a video semantic vector.

[0068] In this embodiment of the invention, after obtaining the surveillance video data, the risk identification system can perform video semantic mining on the surveillance video data to form a video semantic vector. That is, it can mine potential semantic information from the surveillance video data and represent the mining results in vector form, thereby obtaining a video semantic vector.

[0069] Step S130: Perform trajectory semantic mining on the motion trajectory data to form a trajectory semantic vector.

[0070] In this embodiment of the invention, after obtaining the motion trajectory data, the risk identification system can perform trajectory semantic mining on the motion trajectory data to form a trajectory semantic vector. During trajectory semantic mining, different trajectory complexities employ different mining methods. That is, the motion trajectory data can be mined for latent semantic information, and the mining results can be represented in vector form to obtain the trajectory semantic vector.

[0071] Step S140: Aggregate the video semantic vector and the trajectory semantic vector to form a video trajectory aggregate vector.

[0072] In this embodiment of the invention, after obtaining the video semantic vector and the trajectory semantic vector, the risk identification system can aggregate the video semantic vector and the trajectory semantic vector to form a video trajectory aggregation vector. That is, it can aggregate the potential semantic information from both the video and trajectory dimensions, allowing for mutual constraints and ensuring both rich semantic information and accurate semantic information representation. Therefore, the resulting video trajectory aggregation vector has better semantic representation capabilities.

[0073] Step S150: Risk identification is performed based on the video trajectory aggregation vector to obtain the risk identification result.

[0074] In this embodiment of the invention, after obtaining the video trajectory aggregation vector, the risk identification system can perform risk identification based on the video trajectory aggregation vector to obtain a risk identification result. The risk identification result reflects the degree of user behavior risk present in the target area, such as 0-1 or 1-10, with a higher value indicating a higher risk level.

[0075] Furthermore, it should be noted that steps S120-S150 described above can be implemented using a trained neural network model. This neural network model can include an encoder and a decoder. The encoder can be used to perform video semantic mining in step S120, trajectory semantic mining in step S130, and aggregation in step S140, i.e., semantic encoding. The decoder can be used to perform risk identification in step S150, i.e., semantic decoding. Additionally, during training, this neural network model can learn the mapping relationship between sample surveillance video data and sample mobile phone signaling data (corresponding motion trajectory data) and risk labels.

[0076] Based on the above method, since not only video semantic mining but also trajectory semantic mining are performed, the potential semantic information in both video and trajectory dimensions can be mined and aggregated to complete risk identification. This enriches the semantic information on which risk identification is based, thereby improving the reliability of risk identification. Furthermore, because the trajectory semantic mining process also considers the corresponding trajectory complexity, the accuracy of mining potential semantic information in the trajectory dimension can be improved to a certain extent, thereby improving the accuracy of the basis for risk identification and further enhancing the reliability of the risk identification results. This addresses the problem of relatively low reliability in risk identification in existing technologies.

[0077] The specific implementation process of acquiring surveillance video data and motion trajectory data in step S110 of the first part is not limited and can be selected according to the actual situation.

[0078] For example, in one specific implementation, after the corresponding device collects and forms corresponding surveillance video data and mobile phone signaling data, it can first store them in the corresponding database. When risk identification is required, the stored surveillance video data and mobile phone signaling data can be retrieved from the database.

[0079] For example, in another specific implementation, after the corresponding device collects and forms corresponding surveillance video data and mobile phone signaling data, it can send them to the risk identification system, enabling the risk identification system to obtain the surveillance video data and mobile phone signaling data in real time, thereby performing risk identification in real time. Furthermore, it should be noted that if the mobile phone signaling data is GPS positioning data, it can be directly used as motion trajectory data. If the mobile phone signaling data is base station signaling, triangulation can be performed using multiple base station signals to calculate the geographical location of the mobile phone, or the location of the mobile phone can be determined based on the signal strength and the location of the base stations. Alternatively, other existing solutions capable of trajectory determination based on mobile phone signaling data can also be used.

[0080] In the second part, the specific implementation process of performing video semantic mining on the surveillance video data in step S120 is not limited and can be selected according to the actual situation.

[0081] For example, in one specific implementation, each frame of the surveillance video data can be convolved to obtain a corresponding convolution vector. Then, the convolution vectors can be concatenated, added, or averaged to obtain a video semantic vector.

[0082] For example, in another specific implementation, in order to improve the reliability of video semantic mining and enable the obtained video semantic vector to fully represent the potential semantic information related to risk, the above step S120 may further include steps S121, S122 and S123, wherein the details of each step are as follows.

[0083] Step S121: Based on the frame rate of the monitoring video data, determine the target parameters; based on the target parameters, perform frame extraction processing on the monitoring video data to form monitoring video frame-extracted data; and determine the differential video frames between every two adjacent monitoring video frames in the monitoring video frame-extracted data.

[0084] In this embodiment of the invention, combined with Figure 3Based on the frame rate of the monitored video data, a target parameter can be determined. Then, based on the target parameter, frame extraction processing is performed on the monitored video data to form monitored video frame-extracted data. Finally, the difference video frames between every two adjacent monitored video frames in the monitored video frame-extracted data are determined, i.e., the difference between two adjacent monitored video frames is calculated to extract changes in video content. The target parameter and the frame rate are positively correlated; that is, the higher the frame rate, the larger the target parameter. Furthermore, the target parameter can refer to the number of frames between two adjacent monitored video frames in the monitored video frame-extracted data. In other words, because the changes between adjacent monitored video frames are generally smaller at higher frame rates, the need for extracting changes in video content is generally difficult to meet. Therefore, using a larger target parameter for frame extraction allows for more efficient extraction of content changes between monitored video frames.

[0085] Step S122: Semantic mining is performed on each frame of the monitoring video data to form a corresponding semantic vector for each monitoring video frame, and semantic mining is performed on each differential video frame to form a corresponding semantic vector for each differential video frame.

[0086] In this embodiment of the invention, semantic mining can be performed on each frame of the surveillance video data to form a corresponding semantic vector for each surveillance video frame, and semantic mining can also be performed on each frame of the differential video frame to form a corresponding semantic vector for each differential video frame. It should be noted that the methods for semantic mining of the surveillance video frames and the differential video frames can be the same or different. For example, in an alternative implementation, since semantic mining of surveillance video frames tends to focus more on mining global semantic information in the video, while semantic mining of differential video frames tends to focus more on mining changing semantic information in the video, they can be implemented using different convolutional units. The architecture of the convolutional units can be the same, and can include convolutional network layers, pooling network layers, and fully connected network layers.

[0087] Step S123: Semantic fusion is performed on the semantic vector of each monitoring video frame and the semantic vector of each differential video frame to form a video semantic vector.

[0088] In this embodiment of the invention, after obtaining the semantic vector of the monitored video frame and the semantic vector of the differential video frame, semantic fusion can be performed on each of the semantic vectors of the monitored video frame and the differential video frame to form a video semantic vector. For example, each of the semantic vectors of the monitored video frame and the differential video frame can be added, averaged, or otherwise calculated to achieve semantic fusion and obtain the video semantic vector. Alternatively, in other embodiments, each of the semantic vectors of the monitored video frame can be averaged to obtain a first video frame semantic vector, and each of the differential video frame semantic vectors can be averaged to obtain a second video frame semantic vector. Then, the first and second video frame semantic vectors can be concatenated to obtain a corresponding concatenated video frame semantic vector. Finally, self-attention processing is performed on the concatenated video frame semantic vector to achieve the fusion of global video semantic information and video change semantic information to obtain the video semantic vector.

[0089] The third part, the specific implementation process of performing trajectory semantic mining on the motion trajectory data in step S130 is not limited and can be selected according to the actual situation.

[0090] For example, in one specific implementation, for each mobile phone user's motion trajectory data, word embedding (which can be achieved through a trained word embedding model, such as a multilayer perceptron (MLP) / feedforward neural network; for instance, coordinates such as (x,y) or (x,y,z) can be directly used as input and mapped to a fixed-length embedding vector through one or more fully connected layers) can be performed on each coordinate data to obtain a word embedding vector for each coordinate data. Then, the word embedding vectors of each coordinate data can be concatenated to obtain the concatenated vector of the motion trajectory data. Finally, the concatenated vectors of each motion trajectory data can be processed by adding, averaging, or concatenating to obtain the trajectory semantic vector. Specifically, when the trajectory complexity is low, the concatenated vectors of each motion trajectory data can be processed by adding, averaging, or concatenating to obtain the trajectory semantic vector. When the trajectory complexity is high, the concatenated vectors of each motion trajectory data can be processed by adding, averaging, or concatenating to obtain a fused vector. Then, self-attention and other mining techniques can be applied to the fused vector to obtain the trajectory semantic vector.

[0091] For example, in another specific implementation, in order to facilitate the mining of the potential semantic information of the motion trajectory data of each mobile phone user as a whole, so that the obtained trajectory semantic vector can represent the potential semantic information of the association between motion trajectory data, the above step S130 can further include steps S131, S132, S133 and S134, wherein the details of each step are as follows.

[0092] Step S131: Map each trajectory coordinate in each of the motion trajectory data to form a trajectory coordinate mapping parameter corresponding to each trajectory coordinate.

[0093] In this embodiment of the invention, each trajectory coordinate in each of the motion trajectory data can be mapped to form a trajectory coordinate mapping parameter corresponding to each trajectory coordinate. For example, the trajectory coordinates can be serialized, and then the resulting sequence value can be used as the corresponding trajectory coordinate mapping parameter, such as the trajectory coordinate mapping parameter 11 for the first trajectory coordinate, 12 for the second trajectory coordinate, and 13 for the third trajectory coordinate. Alternatively, the obtained sequence value can be normalized to obtain the corresponding trajectory coordinate mapping parameter.

[0094] Step S132: Combine the trajectory coordinate mapping parameters corresponding to each trajectory coordinate in each of the motion trajectory data to form an initial trajectory coordinate matrix; and process the initial trajectory coordinate matrix based on the predetermined target matrix size to form a target trajectory coordinate matrix.

[0095] In this embodiment of the invention, combined with Figure 4 After obtaining the trajectory coordinate mapping parameters, the trajectory coordinate mapping parameters corresponding to each trajectory coordinate in each of the motion trajectory data can be combined to form an initial trajectory coordinate matrix. Then, based on a predetermined target matrix size, the initial trajectory coordinate matrix is ​​processed to form a target trajectory coordinate matrix. The size of the initial trajectory coordinate matrix is ​​less than or equal to the target matrix size, and the size of the target trajectory coordinate matrix is ​​equal to the target matrix size. It should be noted that in the initial trajectory coordinate matrix, trajectory coordinate mapping parameters in the same row correspond to the same motion trajectory data, and trajectory coordinate mapping parameters in the same column correspond to the same time point. Furthermore, if there is no corresponding trajectory coordinate in a certain motion trajectory data at a certain time point, the corresponding trajectory coordinate mapping parameter can be set to 0. And, when processing based on the target matrix size, if it is necessary to increase the size of the initial trajectory coordinate matrix, zeros can be added to the edges to make the size of the resulting target trajectory coordinate matrix equal to the target matrix size.

[0096] Step S133: If the trajectory complexity of the target trajectory coordinate matrix is ​​greater than the predetermined target complexity, then perform multi-scale semantic mining and fusion on the target trajectory coordinate matrix to form a trajectory semantic vector.

[0097] In this embodiment of the invention, after obtaining the target trajectory coordinate matrix, if the trajectory complexity of the target trajectory coordinate matrix is ​​greater than a predetermined target complexity (which can be configured according to actual conditions), then multi-scale semantic mining and fusion are performed on the target trajectory coordinate matrix to form a trajectory semantic vector. The trajectory complexity is used to reflect at least the degree of overlap between different trajectories. For example, it can be measured by the number of overlaps; the greater the number, the higher the trajectory complexity. One overlap refers to two different mobile phone users having the same trajectory coordinates at the same time point. In other words, the higher the trajectory complexity, the higher the probability of conflict between mobile phone users, and correspondingly, the more precise the semantic information required. Therefore, multi-scale semantic mining and fusion can yield a trajectory semantic vector with better semantic representation capabilities.

[0098] Step S134: If the trajectory complexity of the target trajectory coordinate matrix is ​​not greater than the target complexity, then perform single-scale semantic mining on the target trajectory coordinate matrix to form a trajectory semantic vector.

[0099] In this embodiment of the invention, after obtaining the target trajectory coordinate matrix, if the trajectory complexity of the target trajectory coordinate matrix is ​​not greater than the target complexity, then single-scale semantic mining is performed on the target trajectory coordinate matrix to form a trajectory semantic vector. That is, since the lower the trajectory complexity, the lower the probability of conflict between mobile phone users, and correspondingly, the required precision of semantic information is relatively low. Therefore, single-scale semantic mining can be used to make semantic mining more efficient and computationally cheaper. Specifically, single-scale semantic mining can refer to performing convolution, pooling, and fully connected processing on the target trajectory coordinate matrix to achieve deep semantic information mining, and then performing self-attention processing on the mined deep semantic vector to capture potential related semantic information, thereby obtaining the trajectory semantic vector.

[0100] Based on the above implementation method, it is necessary to further explain step S133 above. The specific implementation process of multi-scale semantic mining and fusion of the target trajectory coordinate matrix is ​​not limited. For example, in a specific implementation method, in order to fully mine the potential semantic information of the target trajectory coordinate matrix at different semantic depths, so that the semantic representation accuracy of the formed trajectory semantic vector is higher and a comprehensive representation of multi-depth potential semantic information is achieved, step S133 above may further include steps S133a, S133b, S133c and S133d, wherein the details of each step are as follows.

[0101] Step S133a: The target trajectory coordinate matrix is ​​subjected to feature semantic extraction using multiple different semantic extraction sizes to form a trajectory extraction semantic vector corresponding to each semantic extraction size.

[0102] In this embodiment of the invention, combined with Figure 5 The target trajectory coordinate matrix is ​​subjected to feature semantic extraction using multiple different semantic extraction sizes, forming a trajectory extraction semantic vector corresponding to each extraction size. Feature semantic extraction can be implemented using convolutional network layers, and the semantic extraction size can be 1 / 2, 1 / 4, 1 / 8, 1 / 16, etc., of the size of the target trajectory coordinate matrix. For example, the size of the first trajectory extraction semantic vector is 1 / 2 of the size of the target trajectory coordinate matrix, the size of the second trajectory extraction semantic vector is 1 / 4 of the size of the target trajectory coordinate matrix, the size of the third trajectory extraction semantic vector is 1 / 8 of the size of the target trajectory coordinate matrix, and the size of the fourth trajectory extraction semantic vector is 1 / 16 of the size of the target trajectory coordinate matrix. Alternatively, in other implementations, the target trajectory coordinate matrix can be convolved using a convolutional network layer to obtain a first trajectory extraction semantic vector. Then, the first trajectory extraction semantic vector can be downsampled to obtain a second trajectory extraction semantic vector, and the second trajectory extraction semantic vector can be downsampled to obtain a third trajectory extraction semantic vector, and the third trajectory extraction semantic vector can be downsampled to obtain a fourth trajectory extraction semantic vector.

[0103] Step S133b: For each semantic extraction dimension corresponding to the first dimension category, perform semantic compression operation on the trajectory extraction semantic vector to form a trajectory compressed semantic vector.

[0104] In this embodiment of the invention, after obtaining the trajectory extraction semantic vector, for each trajectory extraction semantic vector belonging to the first size category, a semantic compression operation is performed on the trajectory extraction semantic vector to form a trajectory compressed semantic vector. For example, the larger semantic extraction size can be determined as the first size category, such as the size of the first trajectory extraction semantic vector and the size of the second trajectory extraction semantic vector; and the smaller semantic extraction size can be determined as the second size category, such as the size of the third trajectory extraction semantic vector and the size of the fourth trajectory extraction semantic vector. Based on this, the larger trajectory extraction semantic vector can be further compressed, such as through convolution or pooling, to obtain the corresponding trajectory compressed semantic vector.

[0105] Step S133c: For each semantic extraction dimension corresponding to the second dimension category, perform a semantic expansion operation on the trajectory extraction semantic vector to form a trajectory expansion semantic vector.

[0106] In this embodiment of the invention, after obtaining the trajectory extraction semantic vector, for each trajectory extraction semantic vector belonging to the second size category, a semantic expansion operation is performed on the trajectory extraction semantic vector to form a trajectory expansion semantic vector. Based on this, smaller trajectory extraction semantic vectors can be further expanded, such as through transposed convolution or unpooling, to obtain the corresponding trajectory compression semantic vector. The size of the trajectory expansion semantic vector is equal to the size of the trajectory compression semantic vector.

[0107] Step S133d: Based on the semantic importance corresponding to each of the semantic extraction dimensions, semantically fuse each of the trajectory compressed semantic vectors and each of the trajectory expanded semantic vectors to form a trajectory semantic vector.

[0108] In this embodiment of the invention, after obtaining the trajectory compressed semantic vector and the trajectory expanded semantic vector, each trajectory compressed semantic vector and each trajectory expanded semantic vector are semantically fused based on the semantic importance corresponding to each semantic extraction size to form a trajectory semantic vector. It should be noted that the representational content of semantic vectors obtained based on different semantic extraction sizes is different, and their corresponding importance is also different. Therefore, after performing semantic compression and semantic expansion to achieve size unification, semantic fusion based on the corresponding importance can achieve focused representation of important semantic information, thereby improving the semantic representation capability of the formed trajectory semantic vector. For example, a larger trajectory extraction semantic vector can represent shallow, detailed semantic information, while a smaller trajectory extraction semantic vector can represent deep, abstract semantic information.

[0109] Based on the above implementation method, it should be further explained that the specific implementation process of semantic fusion is not limited. For example, in a specific implementation method, in order to improve the reliability of semantic fusion and make the semantic representation accuracy of the formed trajectory semantic vector higher, the above step S133d may further include step d1, step d2 and step d3, wherein the details of each step are as follows.

[0110] Step d1: Based on the semantic importance corresponding to each semantic extraction size, semantically fuse each trajectory compressed semantic vector and each trajectory expanded semantic vector to form a composite size semantic vector.

[0111] In this embodiment of the invention, combined with Figure 6 Based on the semantic importance corresponding to each of the aforementioned semantic extraction dimensions, the compressed semantic vector and the expanded semantic vector of each trajectory can be semantically fused to form a composite-size semantic vector. As mentioned earlier, since the semantic vectors corresponding to different semantic extraction dimensions focus on different content, and different content has different importance, semantic fusion can be performed based on the corresponding semantic importance to obtain a composite-size semantic vector, thus achieving the initial fusion of semantic information.

[0112] Step d2 involves using multiple dilated convolutional units with different receptive fields to perform depth mining on the composite size semantic vector, forming a trajectory depth semantic vector corresponding to each dilated convolutional unit.

[0113] In this embodiment of the invention, after obtaining the composite-size semantic vector, it can be depth-mined (i.e., dilated convolution is performed, inserting holes (zero values) between the elements of the convolution kernel to expand the receptive field of the convolution, thereby capturing a wider range of contextual information without increasing computational load) through multiple dilated convolution units with different receptive fields, forming a trajectory depth semantic vector corresponding to each dilated convolution unit. It should be noted that performing depth-mining through multiple dilated convolution units with different receptive fields allows for the capture of more semantic information during the depth-mining process. For example, for a smaller receptive field, local details can be perceived; for a slightly larger receptive field, both local details and global information can be considered; and for an even larger receptive field, global context can be perceived.

[0114] Step d3: Based on the semantic importance of each dilated convolutional unit, semantically fuse each trajectory depth semantic vector to form a trajectory semantic vector.

[0115] In this embodiment of the invention, after obtaining each trajectory depth semantic vector, the trajectory depth semantic vectors can be semantically fused based on the semantic importance corresponding to each dilated convolution unit to form a trajectory semantic vector. Based on this, while capturing semantic information from different receptive fields, it is also possible to emphasize different semantic information as needed.

[0116] Based on the above implementation method, it should be further explained that the specific implementation process of semantic fusion is not limited. For example, in a specific implementation method, in order to ensure the accuracy of semantic fusion and to further realize the effective mining of important semantic information during the semantic fusion process, the above step d1 may further include the following (in conjunction with...). Figure 7 ):

[0117] The first step involves linearly mapping each of the trajectory compression semantic vectors and each of the trajectory expansion semantic vectors to form corresponding trajectory linear semantic vectors. Then, gating activation is applied to each of the trajectory linear semantic vectors to form a trajectory gating parameter distribution. The linear mapping can be implemented using a fully connected network layer, and it must maintain the size of the semantic vectors; that is, the size of each trajectory compression semantic vector and each trajectory expansion semantic vector is the same as the size of the corresponding trajectory linear semantic vector. Furthermore, gating activation can map the parameters in the semantic vectors to 0-1, such as using functions like sigmoid, so that the resulting trajectory gating parameter distribution can characterize the importance of the parameters at each position in the corresponding semantic vector, such as 0 representing the lowest importance and 1 representing the highest importance.

[0118] The second step involves adjusting the corresponding trajectory compression semantic vector or trajectory expansion semantic vector for each trajectory gating parameter distribution. (As mentioned earlier, the trajectory gating parameter distribution can characterize the importance of parameters at each position; therefore, during adjustment, bitwise multiplication can be performed to assign corresponding importance to each parameter in the semantic vector.) This forms a trajectory adjustment semantic vector. Furthermore, a gating focus evaluation is performed on the trajectory gating parameter distribution to form a focus evaluation coefficient corresponding to that distribution. This focus evaluation coefficient reflects the uniformity of the distribution of parameters in the trajectory gating parameter distribution. It should be noted that higher uniformity indicates that important and unimportant parameters have not been identified, resulting in a relatively lower focus evaluation coefficient. Conversely, lower uniformity indicates that important and unimportant parameters have been identified, resulting in a relatively higher focus evaluation coefficient. Based on this, the dispersion of each parameter in the trajectory gating parameter distribution can be calculated, and the obtained dispersion can be used as the corresponding focus evaluation coefficient.

[0119] The third step involves adjusting the corresponding focus evaluation coefficient for each semantic extraction size based on a pre-determined importance coefficient corresponding to that semantic extraction size, thereby forming the target weight parameter for that semantic extraction size. For example, the importance coefficient and the focus evaluation coefficient can be multiplied to obtain the target weight parameter. In addition, the importance coefficient can be used as the model parameter of the corresponding neural network model, which can be formed during training and can initially be any value, such as 0, 0.5, 1, etc. Based on this, the obtained target weight parameter can take into account both the learning results of samples and labels and the actual situation of the semantic vector itself.

[0120] Fourthly, based on the target weight parameters corresponding to each semantic extraction size, each trajectory adjustment semantic vector can be weighted and summed to form a composite size semantic vector; for example, the composite size semantic vector = trajectory adjustment semantic vector 1 Target weight parameter 1 + trajectory adjustment semantic vector 2 Target weight parameter 2.

[0121] Based on the above implementation method, it should be further explained that the specific implementation process of semantic fusion is not limited. For example, in a specific implementation method, in order to ensure the accuracy of semantic fusion and to further realize the effective mining of important semantic information during the semantic fusion process, the above step d3 may further include the following (in conjunction with...). Figure 8 ):

[0122] The first step is to perform self-attention processing on each of the trajectory depth semantic vectors to form a corresponding trajectory self-attention semantic vector. Then, based on the distribution of attention parameters during the self-attention processing (i.e., the result of matrix multiplication of the transpose of the query vector and the key vector), attention focus evaluation is performed to obtain the focus evaluation coefficient corresponding to each trajectory self-attention semantic vector. The focus evaluation coefficient is used to reflect the uniformity of the distribution of each parameter in the attention parameter distribution. As mentioned above, the focus evaluation parameter can be determined by the dispersion of each parameter, which reflects the focus effect of the corresponding self-attention processing. The larger the focus evaluation coefficient, the better the focus effect and the higher the importance of the corresponding trajectory self-attention semantic vector.

[0123] The second step involves adjusting the corresponding focus evaluation coefficient for each dilated convolutional unit based on a pre-determined importance coefficient corresponding to the semantic importance of that unit, thus forming the target weight parameter for that unit. For example, the importance coefficient and the focus evaluation coefficient can be multiplied to obtain the target weight parameter. Furthermore, the importance coefficient can serve as a model parameter for the corresponding neural network model, formed during training, and initially can be any value, such as 0, 0.5, or 1. Based on this, the obtained target weight parameter takes into account both the learning results from the samples and labels and the actual situation of the semantic vector itself.

[0124] The third step involves weighting and summing each trajectory self-attention semantic vector based on the target weight parameters corresponding to each dilated convolutional unit to form a trajectory semantic vector; for example, trajectory semantic vector = trajectory self-attention semantic vector 1. The target weight parameters corresponding to dilated convolution unit 1 + trajectory self-attention semantic vector 2 The target weight parameters corresponding to dilated convolution unit 2 + trajectory self-attention semantic vector 3 The target weight parameters corresponding to dilated convolution unit 3.

[0125] In the fourth part, the specific implementation process of aggregating the video semantic vector and the trajectory semantic vector in step S140 is not limited and can be selected according to the actual situation.

[0126] For example, in one specific implementation, the video semantic vector and the trajectory semantic vector can be added, averaged, and concatenated to obtain a video trajectory aggregate vector.

[0127] For example, in another specific implementation, in order to ensure the reliability of semantic aggregation and achieve effective aggregation of semantic information of different dimensions, so that the resulting video trajectory aggregation vector has a high semantic representation capability, the above step S140 may further include steps S141, S142 and S143, wherein the details of each step are as follows.

[0128] Step S141: Attention aggregation is performed on the video semantic vector and the trajectory semantic vector to form a video trajectory attention vector.

[0129] In this embodiment of the invention, attention aggregation can be performed on the video semantic vector and the trajectory semantic vector to form a video trajectory attention vector. The trajectory semantic vector can be mapped to a query vector, and the video semantic vector can be mapped to a key vector and a value vector for cross-attention processing, enabling cross-dimensional fusion of semantic information from both video and trajectory dimensions.

[0130] Step S142: Perform attention aggregation on the video trajectory attention vector and the video semantic vector in multiple stages to form a video trajectory depth vector in multiple stages.

[0131] In this embodiment of the invention, after obtaining the video trajectory attention vector, multiple stages of attention aggregation can be performed on the video trajectory attention vector and the video semantic vector to form multiple stages of video trajectory depth vectors. It should be noted that after attention aggregation based on step S141, considering that the video semantic vector generally contains richer detailed information, further multiple stages of attention aggregation can be performed based on the video semantic vector to fully extract the effective semantic information in the video semantic vector. The object of attention aggregation in each subsequent stage includes the video trajectory depth vector from the previous stage.

[0132] Step S143: Determine the video trajectory aggregation vector based on the video trajectory depth vector of the last stage.

[0133] In this embodiment of the invention, after forming video trajectory depth vectors for multiple stages, a video trajectory aggregation vector can be determined based on the video trajectory depth vector of the last stage. For example, the video trajectory depth vector of the last stage can be determined as the video trajectory aggregation vector.

[0134] Based on the above implementation method, it should be further explained that the specific implementation process of attention aggregation in multiple stages is not limited. For example, in a specific implementation method, in order to fully capture the semantic information in the video semantic vector, the above step S142 may further include the following specific implementation process:

[0135] The first step is to perform attention aggregation on the video trajectory attention vector and the video semantic vector in the first stage of attention aggregation to form the video trajectory depth vector in the first stage; for example, the video trajectory attention vector can be mapped to a query vector, and the video semantic vector can be mapped to a key vector and a value vector for cross-attention processing.

[0136] The second step involves performing attention aggregation on the video trajectory depth vector and the video semantic vector from the previous stage in each subsequent stage to form the video trajectory depth vector for the current stage. For example, the video trajectory depth vector from the first stage can be mapped to a query vector, and the video semantic vector can be mapped to a key vector and a value vector for cross-attention processing to obtain the video trajectory depth vector for the second stage.

[0137] The fifth part, the specific implementation process of risk identification based on the video trajectory aggregation vector in step S150 is not limited and can be selected according to the actual situation.

[0138] For example, in one specific implementation, the video trajectory aggregation vector can be fully connected to obtain a corresponding fully connected video trajectory vector, wherein the size of the fully connected video trajectory vector can be 1. 1. This includes a vector parameter. Then, the fully connected vector of the video trajectory can be linearly mapped or identity-mapped to obtain a target parameter, which can be used to characterize the degree of user behavior risk. For example, if neither the surveillance video data nor the motion trajectory data indicates users approaching each other, the target parameter can be 0, indicating no behavioral risk. If both the surveillance video data and the motion trajectory data indicate a large number of users approaching, touching, pushing, etc., the target parameter can be 0.8, etc.

[0139] In summary, the risk identification method and system based on video data and mobile signaling analysis provided by this invention first acquires surveillance video data and mobile signaling data generated in a target area within a target time interval, and determines the motion trajectory data of each mobile phone user based on the mobile signaling data; secondly, it performs video semantic mining on the surveillance video data to form video semantic vectors; then, it performs trajectory semantic mining on the motion trajectory data to form trajectory semantic vectors, wherein different trajectory complexities employ different mining methods during trajectory semantic mining; further, it aggregates the video semantic vectors and trajectory semantic vectors to form a video trajectory aggregation vector; finally, it performs risk identification based on the video trajectory aggregation vector to obtain the risk identification result. Based on the above method, because it performs not only video semantic mining but also trajectory semantic mining, the potential semantic information of both video and trajectory dimensions can be mined and aggregated to complete risk identification, making the semantic information on which risk identification is based richer, thereby improving the reliability of risk identification. Furthermore, since trajectory complexity is also considered during trajectory semantic mining, the accuracy of mining potential semantic information in the trajectory dimension can be improved to a certain extent, thereby improving the accuracy of the basis for risk identification, further improving the reliability of risk identification results, and thus improving the problem of relatively low reliability of risk identification in existing technologies.

[0140] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0141] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0142] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0143] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A risk identification method based on video data and mobile phone signaling analysis, characterized in that, include: Acquire surveillance video data and mobile phone signaling data generated in the target area within the target time interval, and determine the motion trajectory data of each mobile phone user based on the mobile phone signaling data; Perform video semantic mining on the surveillance video data to form video semantic vectors; The motion trajectory data is subjected to trajectory semantic mining to form a trajectory semantic vector, including: mapping each trajectory coordinate in each of the motion trajectory data to form a trajectory coordinate mapping parameter corresponding to each trajectory coordinate; combining the trajectory coordinate mapping parameters corresponding to each trajectory coordinate in each of the motion trajectory data to form an initial trajectory coordinate matrix; and processing the initial trajectory coordinate matrix based on a predetermined target matrix size to form a target trajectory coordinate matrix, wherein the size of the initial trajectory coordinate matrix is ​​less than or equal to the size of the target matrix, and the size of the target trajectory coordinate matrix is ​​equal to the size of the target matrix; if the trajectory complexity of the target trajectory coordinate matrix is ​​greater than a predetermined target complexity, then multi-scale semantic mining and fusion are performed on the target trajectory coordinate matrix to form a trajectory semantic vector, wherein the trajectory complexity is at least used to reflect the degree of overlap between different trajectories; if the trajectory complexity of the target trajectory coordinate matrix is ​​not greater than the target complexity, then single-scale semantic mining is performed on the target trajectory coordinate matrix to form a trajectory semantic vector. The video semantic vector and the trajectory semantic vector are aggregated to form a video trajectory aggregate vector; Risk identification is performed based on the video trajectory aggregation vector to obtain risk identification results, wherein the risk identification results are used to reflect the degree of user behavior risk existing in the target area.

2. The risk identification method based on video data and mobile signaling analysis as described in claim 1, characterized in that, The step of performing multi-scale semantic mining and fusion on the target trajectory coordinate matrix to form a trajectory semantic vector if the trajectory complexity of the target trajectory coordinate matrix is ​​greater than a predetermined target complexity includes: The target trajectory coordinate matrix is ​​subjected to feature semantic extraction using multiple different semantic extraction sizes to form a trajectory extraction semantic vector corresponding to each semantic extraction size. For each semantic extraction dimension belonging to the first dimension category, the trajectory extraction semantic vector is subjected to semantic compression to form a trajectory compressed semantic vector. For each semantic extraction size belonging to the second size category, the trajectory extraction semantic vector is subjected to a semantic expansion operation to form a trajectory expansion semantic vector, wherein the size of the trajectory expansion semantic vector is equal to the size of the trajectory compression semantic vector. Based on the semantic importance corresponding to each of the semantic extraction dimensions, each trajectory compressed semantic vector and each trajectory expanded semantic vector are semantically fused to form a trajectory semantic vector.

3. The risk identification method based on video data and mobile signaling analysis as described in claim 2, characterized in that, The step of semantically fusing each trajectory compressed semantic vector and each trajectory expanded semantic vector based on the semantic importance corresponding to each semantic extraction size to form a trajectory semantic vector includes: Based on the semantic importance corresponding to each semantic extraction size, each trajectory compressed semantic vector and each trajectory expanded semantic vector are semantically fused to form a composite size semantic vector. The composite size semantic vector is deeply mined by using multiple dilated convolutional units with different receptive fields to form a trajectory depth semantic vector corresponding to each dilated convolutional unit. Based on the semantic importance of each dilated convolutional unit, the trajectory depth semantic vector is semantically fused to form a trajectory semantic vector.

4. The risk identification method based on video data and mobile signaling analysis as described in claim 3, characterized in that, The step of semantically fusing each trajectory compressed semantic vector and each trajectory expanded semantic vector based on the semantic importance corresponding to each semantic extraction size to form a composite size semantic vector includes: Linear mapping is performed on each of the trajectory compression semantic vectors and each of the trajectory expansion semantic vectors to form corresponding trajectory linear semantic vectors; and gating activation is performed on each of the trajectory linear semantic vectors to form a trajectory gating parameter distribution. For each of the trajectory gating parameter distributions, the corresponding trajectory compression semantic vector or trajectory expansion semantic vector is adjusted based on the trajectory gating parameter distribution to form a trajectory adjustment semantic vector. Furthermore, a gating focus evaluation is performed on the trajectory gating parameter distribution to form a focus evaluation coefficient corresponding to the trajectory gating parameter distribution. The focus evaluation coefficient is used to reflect the uniformity of the distribution of each parameter in the trajectory gating parameter distribution. For each semantic extraction size, the corresponding focus evaluation coefficient is adjusted based on the importance coefficient that is predetermined for the semantic importance corresponding to that semantic extraction size, to form the target weight parameter corresponding to that semantic extraction size; Based on the target weight parameters corresponding to each semantic extraction size, each trajectory adjustment semantic vector is weighted and summed to form a composite size semantic vector.

5. The risk identification method based on video data and mobile signaling analysis as described in claim 3, characterized in that, The step of semantically fusing each trajectory depth semantic vector based on the semantic importance corresponding to each dilated convolutional unit to form a trajectory semantic vector includes: Each trajectory depth semantic vector is subjected to self-attention processing to form a corresponding trajectory self-attention semantic vector. Attention focus evaluation is performed based on the attention parameter distribution during the self-attention processing to obtain the focus evaluation coefficient corresponding to each trajectory self-attention semantic vector. The focus evaluation coefficient is used to reflect the uniformity of the distribution of each parameter in the attention parameter distribution. For each of the dilated convolutional units, the corresponding focus evaluation coefficients are adjusted based on the importance coefficients that are predetermined for the semantic importance of the dilated convolutional unit, to form the target weight parameters corresponding to the dilated convolutional unit. Based on the target weight parameters corresponding to each dilated convolutional unit, each trajectory self-attention semantic vector is weighted and summed to form a trajectory semantic vector.

6. The risk identification method based on video data and mobile signaling analysis as described in claim 1, characterized in that, The step of aggregating the video semantic vector and the trajectory semantic vector to form a video trajectory aggregate vector includes: Attention aggregation is performed on the video semantic vector and the trajectory semantic vector to form a video trajectory attention vector; The video trajectory attention vector and the video semantic vector are subjected to multiple stages of attention aggregation to form multiple stages of video trajectory depth vectors, wherein the object of attention aggregation in a later stage includes the video trajectory depth vector of the previous stage. Based on the video trajectory depth vector of the last stage, the video trajectory aggregation vector is determined.

7. The risk identification method based on video data and mobile signaling analysis as described in claim 6, characterized in that, The step of performing multi-stage attention aggregation on the video trajectory attention vector and the video semantic vector to form a multi-stage video trajectory depth vector includes: In the first stage of attention aggregation, the video trajectory attention vector and the video semantic vector are aggregated to form the video trajectory depth vector of the first stage. In each subsequent attention aggregation stage, attention aggregation is performed on the video trajectory depth vector and the video semantic vector from the previous stage to form the video trajectory depth vector for the current stage.

8. The risk identification method based on video data and mobile signaling analysis as described in any one of claims 1-7, characterized in that, The step of performing video semantic mining on the surveillance video data to form video semantic vectors includes: Based on the frame rate of the monitoring video data, a target parameter is determined, and based on the target parameter, the monitoring video data is processed by frame extraction to form monitoring video frame-extracted data, and the differential video frames between every two adjacent monitoring video frames in the monitoring video frame-extracted data are determined respectively, wherein the target parameter and the frame rate have a positive correlation. Semantic mining is performed on each frame of the surveillance video data to form a semantic vector for each surveillance video frame. Semantic mining is also performed on each frame of the differential video frame to form a semantic vector for each differential video frame. Semantic fusion is performed on the semantic vector of each of the monitored video frames and the semantic vector of each of the differential video frames to form a video semantic vector.

9. A risk identification system based on video data and mobile phone signaling analysis, characterized in that, The device includes a processor and a memory, the memory being used to store a computer program, and the processor being used to execute the computer program to implement the risk identification method based on video data and mobile signaling analysis as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Identifying method and device for illegal use of mobile phone, medium and equipment

    CN115460548A

  • Airport video data real-time analysis system

    CN120976826A