An electronic evidence identification method and system
By cutting and tampering judgment on video stream information, combined with neural network screening and feature extraction, the problem of low accuracy and efficiency of video data in identifying illegal activities is solved, and more efficient and reliable electronic evidence processing and illegal activity identification are achieved.
Patent Information
- Application Number
- CN202510050039.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The complexity of video data scenarios leads to low accuracy and efficiency in identifying illegal activities.
By obtaining video stream information and performing shot cutting to obtain shot information, it is determined whether it has been tampered with. If it has not been tampered with, screening and feature extraction are performed, and a neural network is used to identify illegal activities.
It improves the credibility of the electronic evidence processing process and the efficiency of identifying illegal activities, ensures the reliability and accuracy of analysis, and enhances the fairness and accuracy of law enforcement.
Smart Images

Figure CN119992450B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for electronic evidence recognition based on a neural network. Background Art
[0002] With the rapid development of information technology and the widespread adoption of digital devices and network applications, vast quantities of electronic data are being generated in all aspects of social production and life. From everyday social chat logs and emails to electronic contracts and transaction records in commercial activities, and even video footage in surveillance systems, electronic data, with its convenient storage, transmission, and replication capabilities, has gradually become a vital source of evidence. The role of electronic evidence is becoming increasingly prominent in numerous fields, including judicial practice, commercial dispute resolution, and law enforcement, and has even become a key basis for convictions in some cases. As a form of electronic evidence, video information can clearly record the time, location, and behavior of a suspect. Therefore, video data is increasingly being used as electronic evidence in the field of illegal behavior identification. However, due to the complex nature of video data, its accuracy and efficiency in identifying illegal behavior suffer from low accuracy and efficiency. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for electronic evidence identification to improve the above problems.
[0004] In order to achieve the above objectives, the embodiments of the present application provide the following technical solutions:
[0005] In one aspect, an embodiment of the present application provides a method for identifying electronic evidence, the method comprising:
[0006] Acquiring electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information;
[0007] Performing shot cutting on the video stream information to obtain at least one shot information, where the shot information includes at least two consecutive image frames;
[0008] determining, based on the lens information, whether the electronic evidence information to be identified has been tampered with, to obtain a first determination result;
[0009] When the first judgment result is that the electronic evidence information to be identified has not been tampered with, screening the electronic evidence information to be identified to obtain a screening result, wherein the screening result includes an image frame in which the target action exists;
[0010] The illegal behavior of the target object is identified according to the screening result to obtain an identification result.
[0011] In a second aspect, an embodiment of the present application provides an electronic evidence recognition system based on a neural network, the system comprising:
[0012] An acquisition module, configured to acquire electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information;
[0013] A first processing module is configured to perform shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two consecutive image frames;
[0014] a judgment module, configured to judge whether the electronic evidence information to be identified has been tampered with based on the lens information, and obtain a first judgment result;
[0015] a second processing module configured to, when the first judgment result is that the electronic evidence information to be identified has not been tampered with, filter the electronic evidence information to be identified to obtain a filtering result, wherein the filtering result includes an image frame in which a target action is present;
[0016] The third processing module is used to identify the illegal behavior of the target object according to the screening result to obtain an identification result.
[0017] In a third aspect, embodiments of the present application provide a neural network-based electronic evidence identification device, comprising a memory and a processor. The memory is configured to store a computer program; the processor is configured to implement the steps of the neural network-based electronic evidence identification method when executing the computer program.
[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned neural network-based electronic evidence identification method.
[0019] The beneficial effects of the present invention are:
[0020] The present invention obtains multiple shot information by performing lens cutting on the electronic evidence information to be identified, and then determines whether each shot information has been tampered with, so as to ensure that the subsequent analysis and judgment based on the evidence are built on a reliable foundation, thereby improving the credibility of the entire electronic evidence processing process. When the electronic evidence information to be identified has not been tampered with, the electronic evidence information to be identified is screened to obtain the screening results, and finally the illegal behavior of the target object is identified based on the screening results, thereby improving the efficiency of electronic evidence in identifying illegal behavior.
[0021] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 Schematic diagram of the electronic evidence identification method in the present invention.
[0024] Figure 2 Schematic diagram of the topology of the electronic evidence identification system in the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.
[0027] Example 1:
[0028] This embodiment provides an electronic evidence identification method based on a neural network. It can be understood that in this embodiment, a scenario can be set up, for example: processing the video stream data collected by a surveillance camera to determine whether there is any illegal behavior in the video stream data.
[0029] See also Figure 1 , the figure shows that the method includes step S1, step S2, step S3, step S4 and step S5.
[0030] Step S1: obtaining electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information;
[0031] Step S2: performing shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two consecutive image frames;
[0032] Among them, the video stream is composed of frames, shots, and scenes, which are continuously stacked and spliced in chronological order. The video stream information is cut into shots to obtain shot information, and continuous image frames are analyzed as a whole, which provides a more detailed and accurate basic unit for subsequent judgment and analysis, so that the processing of video content is not based on scattered images, but on coherent shots, which can better reflect the actual content of the video and scene changes.
[0033] Step S3: determining whether the electronic evidence information to be identified has been tampered with based on the lens information, and obtaining a first determination result;
[0034] In the application of electronic evidence, the authenticity and integrity of the evidence are of vital importance. This step uses lens information to determine whether the electronic evidence information to be identified has been tampered with. It can promptly detect human modifications to video evidence and effectively exclude tampered video evidence, ensuring that subsequent evidence-based analysis and judgment are based on a reliable foundation, thereby improving the credibility of the entire electronic evidence processing process.
[0035] Step S4: When the first judgment result is that the electronic evidence information to be identified has not been tampered with, the electronic evidence information to be identified is screened to obtain a screening result, wherein the screening result includes an image frame in which the target action exists;
[0036] Furthermore, the step S4 includes step S41, step S42 and step S43:
[0037] Step S41: sending the image frames in the electronic evidence information to be identified to a human posture estimation model to obtain skeleton position information corresponding to each image frame;
[0038] Step S42: extracting features from the skeleton position information corresponding to each image frame to obtain feature sequence information;
[0039] Step S43: Send the feature sequence information to the spatiotemporal feature long short-term memory network to obtain a screening result.
[0040] In this step, the temporal dependency and spatial dependency between human joints can be modeled simultaneously based on the spatiotemporal feature long short-term memory network to complete the detection of the target action, thereby being able to quickly locate the key images related to the target action from a large amount of video information, avoiding indiscriminate analysis of the entire video, greatly reducing the amount of data processing, improving work efficiency, and highlighting the key analysis content.
[0041] Step S5: Identify the illegal behavior of the target object based on the screening results to obtain an identification result, so that the determination of the illegal behavior is more accurate and objective, which helps to improve the fairness and accuracy of law enforcement.
[0042] Example 2:
[0043] The only difference between this embodiment and embodiment 1 is that step S3, determining whether the electronic evidence information to be identified has been tampered with based on the lens information, and obtaining a first judgment result, includes steps S31 to S33:
[0044] Step S31: determining the key frames included in each shot information according to the shot information, which specifically includes the following steps:
[0045] Step S311, acquiring two adjacent image frames;
[0046] Step S312: Calculate the mean, variance, and covariance in the local window of two adjacent image frames to obtain a calculation result;
[0047] Step S313: determining the similarity between two adjacent image frames according to the calculation result to obtain similarity information;
[0048] In this step, the specific calculation formula of similarity is:
[0049]
[0050] In the above formula, S represents similarity; α1 and α2 represent the mean values in the local windows of two adjacent image frames respectively; and Respectively represent the variance in the local window of two adjacent image frames; σ 12 represents the covariance within the local window of two adjacent image frames, and c1 and c2 represent constants.
[0051] Step S314: Determine the key frame based on the similarity information between each two adjacent image frames. In this step, by calculating the similarity between each two adjacent image frames, the frame sequence can be used as the horizontal coordinate and the similarity as the vertical coordinate to draw a similarity curve. The similarity curve is smoothed and denoised. The processed similarity curve can measure whether there are large differences between adjacent frames to extract key frames.
[0052] Step S32: pre-processing the key frame to obtain a pre-processed key frame, wherein the pre-processing includes enhancing the key frame to improve image quality and enrich information.
[0053] Step S33: judging whether the electronic evidence information to be identified has been tampered with based on the preprocessed key frame, and obtaining a first judgment result.
[0054] When determining whether electronic evidence has been tampered with, operations such as editing and splicing will leave more obvious traces on the key frames. By capturing the key frames included in each shot information and then preprocessing the key frames, the characteristics of the key frames are further optimized, making subsequent tampering judgments based on the preprocessed key frames more accurate, and improving the ability to identify tampering features, thereby more accurately judging the authenticity of electronic evidence.
[0055] Specifically, the step S33 includes steps S331 to S335:
[0056] Step S331: extracting feature information corresponding to the pre-processed key frame in the spatial domain to obtain first feature information;
[0057] In this step, the preprocessed key frame can be sent to the multi-scale fusion network for processing multiple times to obtain the first feature information, wherein the multi-scale fusion network processing includes first increasing the dimension through the 1×1 convolution layer, and then reducing the computational complexity through the 3×3 DW convolution, and then using the dilated convolution with different dilation factors to capture the multi-scale tampering detail information and splice it on the channel, and then learning the importance of each scale feature channel through the global average pooling and two fully connected layers to obtain the local discontinuity of the preprocessed key frame, and then 1×1 convolution is used for dimensionality reduction, and then 1×1 dimensionality reduction PW convolution and random inactivation layer are used to prevent overfitting to obtain forged features, and finally added to itself and formed through the channel-spatial attention module to form the final multi-scale fusion feature, i.e., the first feature information.
[0058] Step S332: Extract feature information corresponding to the preprocessed key frame in the frequency domain to obtain second feature information. For example, in this step, a frequency domain image can be obtained through a high-frequency noise stream, and then the frequency domain image can be sent to a multi-scale fusion network multiple times to obtain high-frequency noise information, i.e., the second feature information.
[0059] Step S333: concatenate the first characteristic information and the second characteristic information to obtain concatenated characteristic information;
[0060] Step S334: sending the spliced feature information to the attention network to obtain weighted feature information;
[0061] In this step, the channel attention mechanism can be used to calculate the channel dimension information, and then the spatial dimension information can be calculated using the spatial attention mechanism and residual connection can be performed with the spliced feature information to obtain the interactive features. Finally, the original feature map and the weighted feature map are added to obtain the degree of influence of the first feature information and the second feature information on the tampered information in their respective channels.
[0062] Step S335: Determine whether the electronic evidence information to be identified has been tampered with based on the weighted feature information to obtain a first judgment result.
[0063] When the attention mechanism is used to fuse the first feature information in the spatial domain and the second feature information in the frequency domain, the feature streams of the two different modalities can interact and promote learning instead of keeping them independent, thereby capturing the inconsistency between the tampered area and the real area, and accurately judging whether the electronic evidence information to be identified has been tampered with.
[0064] Example 3:
[0065] The difference between this embodiment and embodiment 1 or 2 is that step S5, identifying the illegal behavior of the target object according to the screening result to obtain the identification result, includes steps S51 to S54:
[0066] Step S51: sending the screening result to a spatial feature extraction network to obtain third feature information;
[0067] In surveillance scenarios with a wide field of view, small human targets, clusters of people, and occlusions make it difficult for traditional methods to grasp the spatial relationships between human subjects. Spatial feature extraction networks focus on the spatial structure of images, deeply analyzing and extracting the spatial position, shape, and relative positional relationships of target objects in the screening results. This helps capture the underlying relationships between human subjects, avoiding misjudgments of target objects due to loss or confusion of local details, and thus providing a more accurate spatial information foundation for subsequent illegal behavior identification.
[0068] Step S52: sending the electronic evidence information to be identified to a spatiotemporal feature extraction network to obtain fourth feature information;
[0069] In surveillance videos, the behavior of target objects changes over time, and there are dynamic interactions between the object and its environment, and between each other. The spatiotemporal feature extraction network fully considers the temporal dimension of the video, extracting not only spatial features from each frame but also capturing the target object's motion trajectory, behavioral changes, and interaction with its surroundings at different moments. This helps to better understand the target object's behavioral patterns and their relationship with the surrounding environment in complex surveillance scenarios, overcoming the problem that traditional methods often break down the underlying relationships between the object and its environment, and between each other.
[0070] Step S53: sending the third feature information and the fourth feature information to a feature fusion attention network to obtain fused feature information;
[0071] The feature fusion attention network can deeply fuse feature information from different channels and automatically learn the importance of different features using the attention mechanism. In complex monitoring scenarios, there is a large amount of information, some of which is crucial for identifying illegal behavior, while some may be interference information. The attention mechanism can focus on key features, give higher weights to these important features, and suppress irrelevant or minor features. In this way, the most representative and discriminative information can be extracted from the fused feature information, further improving the accuracy of identifying illegal behavior of target objects in complex scenarios.
[0072] Step S54: Identify the illegal behavior of the target object based on the fused feature information to obtain an identification result.
[0073] Specifically, step S54 includes steps S541 to S545:
[0074] Step S541: Determine whether a target object corresponding to the target action exists in the screening results;
[0075] Step S542: segmenting the face image of the target object to obtain face image information;
[0076] Step S543: Determine whether the facial image information is blocked, and obtain a second determination result;
[0077] In most scenarios, the face may be obstructed by objects such as sunglasses, scarves, and masks, which may make it impossible to accurately identify the target object's emotions, thereby reducing the accuracy of illegal behavior identification. Therefore, when identifying illegal behavior, it is first necessary to determine whether the target object's facial image information is obstructed.
[0078] Step S544: Identify the emotion of the target object according to the second judgment result to obtain an emotion recognition result;
[0079] In this step, when the second judgment result is that there is occlusion in the facial image of the target object, the occluded part in the facial image information is determined to obtain the occluded image information; the occluded image information and the facial image information are sent to a partial convolutional generative adversarial network to obtain a de-occluded facial image; the emotion of the target object is identified based on the de-occluded facial image to obtain an emotion recognition result, and the emotion of the target object is accurately determined by repairing the occluded facial image, thereby improving the applicability of the present invention and making it applicable to most scenarios.
[0080] Step S545: Identify the illegal behavior of the target object based on the emotion recognition result and the fused feature information.
[0081] Identifying illegal behaviors based solely on the target actions of a single modal target object may result in misidentification. Therefore, in this embodiment, identifying the illegal behaviors of the target object in combination with the emotion recognition results of the target object can effectively reduce the misidentification rate of illegal behaviors.
[0082] Example 4:
[0083] like Figure 2 As shown, this embodiment provides a neural network-based electronic evidence identification system, which includes an acquisition module 901, a first processing module 902, a judgment module 903, a second processing module 904, and a third processing module 905, which specifically include:
[0084] An acquisition module 901 is configured to acquire electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information;
[0085] A first processing module 902 is configured to perform shot segmentation on the video stream information to obtain at least one shot information, where the shot information includes at least two consecutive image frames;
[0086] A judgment module 903 is configured to judge whether the electronic evidence information to be identified has been tampered with based on the lens information, and obtain a first judgment result;
[0087] A second processing module 904 is configured to, when the first judgment result is that the electronic evidence information to be identified has not been tampered with, filter the electronic evidence information to be identified to obtain a filtering result, wherein the filtering result includes an image frame containing a target action;
[0088] The third processing module 905 is used to identify the illegal behavior of the target object according to the screening result to obtain an identification result.
[0089] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
[0090] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for identifying electronic evidence, characterized in that: include: Acquiring electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information; Performing shot cutting on the video stream information to obtain at least one shot information, where the shot information includes at least two consecutive image frames; determining, based on the lens information, whether the electronic evidence information to be identified has been tampered with, to obtain a first determination result; When the first judgment result is that the electronic evidence information to be identified has not been tampered with, screening the electronic evidence information to be identified to obtain a screening result, wherein the screening result includes an image frame in which the target action exists; Identify the illegal behavior of the target object according to the screening result to obtain an identification result; Wherein, determining whether the electronic evidence information to be identified has been tampered with according to the lens information to obtain a first judgment result includes: Determining the key frames included in each shot information according to the shot information; Preprocessing the key frame to obtain a preprocessed key frame; Determining whether the electronic evidence information to be identified has been tampered with based on the preprocessed key frame to obtain a first determination result; Wherein, determining whether the electronic evidence information to be identified has been tampered with based on the preprocessed key frame to obtain a first judgment result includes: Extracting feature information corresponding to the preprocessed key frame in the spatial domain to obtain first feature information; Extracting feature information corresponding to the preprocessed key frame in the frequency domain to obtain second feature information; concatenating the first characteristic information and the second characteristic information to obtain concatenated characteristic information; Sending the spliced feature information to the attention network to obtain weighted feature information; Determining whether the electronic evidence information to be identified has been tampered with based on the weighted feature information to obtain a first determination result; The illegal behavior of the target object is identified based on the screening results to obtain the identification results, including: Sending the screening results to a spatial feature extraction network to obtain third feature information; Sending the electronic evidence information to be identified to the spatiotemporal feature extraction network to obtain fourth feature information; Sending the third feature information and the fourth feature information to a feature fusion attention network to obtain fused feature information; The illegal behavior of the target object is identified based on the fused feature information to obtain an identification result.
2. The electronic evidence identification method according to claim 1, characterized in that: Determining a key frame included in each shot information according to the shot information includes: Acquire two adjacent image frames; Calculate the mean, variance and covariance in the local window of two adjacent image frames to obtain the calculation results; Determining the similarity between two adjacent image frames according to the calculation result to obtain similarity information; The key frame is determined based on the similarity information between each two adjacent image frames.
3. The electronic evidence identification method according to claim 1, characterized in that: Identify the illegal behavior of the target object based on the fused feature information to obtain an identification result, including: Determine whether a target object corresponding to the target action exists in the screening results; Segmenting the face image of the target object to obtain face image information; Determining whether the facial image information is blocked, and obtaining a second determination result; Identify the emotion of the target object according to the second judgment result to obtain an emotion recognition result; The illegal behavior of the target object is identified based on the emotion recognition result and the fused feature information.
4. The electronic evidence identification method according to claim 3, characterized in that: Identifying the emotion of the target object according to the second judgment result to obtain an emotion recognition result includes: When the second judgment result is that the facial image of the target object is occluded, determining the occluded portion in the facial image information to obtain occluded image information; Sending the occluded image information and the facial image information to a partial convolutional generative adversarial network to obtain a deoccluded facial image; The emotion of the target object is recognized based on the de-occluded facial image to obtain an emotion recognition result.
5. The electronic evidence identification method according to claim 2, characterized in that: The specific calculation method of similarity is: In the above formula, S represents similarity; α1 and α2 represent the mean values in the local windows of two adjacent image frames respectively; and Respectively represent the variance in the local window of two adjacent image frames; σ 12 represents the covariance within the local window of two adjacent image frames, and c1 and c2 represent constants.
6. The electronic evidence identification method according to claim 1, characterized in that: Screening the electronic evidence information to be identified to obtain screening results, including: Sending the image frames in the electronic evidence information to be identified to a human posture estimation model to obtain skeleton position information corresponding to each image frame; Extract features from the skeleton position information corresponding to each image frame to obtain feature sequence information; The feature sequence information is sent to a spatiotemporal feature long short-term memory network to obtain a screening result.
7. A neural network-based electronic evidence recognition system, characterized in that: include: An acquisition module, configured to acquire electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information; A first processing module is configured to perform shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two consecutive image frames; a judgment module, configured to judge whether the electronic evidence information to be identified has been tampered with based on the lens information, and obtain a first judgment result; a second processing module configured to, when the first judgment result is that the electronic evidence information to be identified has not been tampered with, filter the electronic evidence information to be identified to obtain a filtering result, wherein the filtering result includes an image frame in which a target action is present; A third processing module is used to identify the illegal behavior of the target object according to the screening result to obtain an identification result; Wherein, the judgment module includes: Determining the key frames included in each shot information according to the shot information; Preprocessing the key frame to obtain a preprocessed key frame; Determining whether the electronic evidence information to be identified has been tampered with based on the preprocessed key frame to obtain a first determination result; Wherein, determining whether the electronic evidence information to be identified has been tampered with based on the preprocessed key frame to obtain a first judgment result includes: Extracting feature information corresponding to the preprocessed key frame in the spatial domain to obtain first feature information; Extracting feature information corresponding to the preprocessed key frame in the frequency domain to obtain second feature information; concatenating the first characteristic information and the second characteristic information to obtain concatenated characteristic information; Sending the spliced feature information to the attention network to obtain weighted feature information; Determining whether the electronic evidence information to be identified has been tampered with based on the weighted feature information to obtain a first determination result; Wherein, the third processing module includes: Sending the screening results to a spatial feature extraction network to obtain third feature information; Sending the electronic evidence information to be identified to the spatiotemporal feature extraction network to obtain fourth feature information; Sending the third feature information and the fourth feature information to a feature fusion attention network to obtain fused feature information; The illegal behavior of the target object is identified based on the fused feature information to obtain an identification result.
Citation Information
Patent Citations
Motor vehicle violation shooting system and method
CN106408952A
Tampered video identification method and device, electronic equipment and storage medium
CN118799772A