Electronic evidence identification method and system

By performing lens cutting and tampering judgment on video data, untampered information is screened out and illegal behavior is used to identify illegal behaviors, the problem of low accuracy and efficiency of video data when identifying illegal behaviors is solved, and the credibility and recognition efficiency of electronic evidence processing are improved.

CN119992450AActive Publication Date: 2025-05-13WUHAN PKU HIGH-TECH SOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510050039.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13
Estimated Expiration
2045-01-13

Smart Images

  • Figure CN119992450A_ABST
    Figure CN119992450A_ABST
Patent Text Reader

Abstract

The invention relates to an electronic evidence identification method and system, and the method comprises the steps: obtaining to-be-identified electronic evidence information, and the to-be-identified electronic evidence information comprises video stream information; performing shot cutting on the video stream information to obtain at least one piece of shot information; judging whether the to-be-identified electronic evidence information is tampered or not according to the lens information to obtain a first judgment result; when the first judgment result is that the to-be-identified electronic evidence information is not tampered, screening the to-be-identified electronic evidence information to obtain a screening result, the screening result comprising an image frame with a target action; according to the method, a series of processes from authenticity judgment of the electronic evidence to illegal behavior recognition are realized, the efficiency of evidence analysis is greatly improved, and a large amount of manpower, material resources and time cost are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for electronic evidence recognition based on a neural network. Background Art

[0002] With the rapid development of information technology, the widespread popularity of digital devices and network applications, a large amount of electronic data has been generated in all aspects of social production and life. From daily social chat records and e-mails to electronic contracts and transaction records in business activities, and video data in monitoring systems, electronic data has gradually become an important source of evidence with its convenient storage, transmission and copying characteristics. In many fields such as judicial practice, commercial dispute resolution, and law enforcement actions, the role of electronic evidence has become increasingly prominent, and even in some cases has become a key basis for conviction. Video information, as a type of electronic evidence, can clearly record the time, place and behavior of the suspect's crime. Therefore, video data is increasingly used as electronic evidence in the field of illegal behavior identification. However, due to the complexity of video data scenarios, it has low accuracy and low efficiency in identifying illegal behaviors. Summary of the invention

[0003] The purpose of the present invention is to provide an electronic evidence identification method and system to improve the above problems.

[0004] In order to achieve the above objectives, the present application provides the following technical solutions:

[0005] On the one hand, an embodiment of the present application provides a method for identifying electronic evidence, the method comprising:

[0006] Acquiring electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information;

[0007] Performing shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two continuous image frames;

[0008] Determining whether the electronic evidence information to be identified has been tampered with according to the lens information, and obtaining a first determination result;

[0009] When the first judgment result is that the electronic evidence information to be identified has not been tampered with, screening the electronic evidence information to be identified to obtain a screening result, wherein the screening result includes an image frame in which the target action exists;

[0010] The illegal behavior of the target object is identified according to the screening result to obtain an identification result.

[0011] In a second aspect, an embodiment of the present application provides an electronic evidence recognition system based on a neural network, the system comprising:

[0012] An acquisition module, used for acquiring electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information;

[0013] A first processing module, configured to perform shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two continuous image frames;

[0014] A judgment module, used for judging whether the electronic evidence information to be identified has been tampered with according to the lens information, and obtaining a first judgment result;

[0015] A second processing module is configured to, when the first judgment result is that the electronic evidence information to be identified has not been tampered with, screen the electronic evidence information to be identified to obtain a screening result, wherein the screening result includes an image frame in which a target action exists;

[0016] The third processing module is used to identify the illegal behavior of the target object according to the screening result to obtain an identification result.

[0017] In a third aspect, an embodiment of the present application provides a neural network-based electronic evidence identification device, the device comprising a memory and a processor. The memory is used to store a computer program; the processor is used to implement the steps of the neural network-based electronic evidence identification method when executing the computer program.

[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-mentioned neural network-based electronic evidence identification method are implemented.

[0019] The beneficial effects of the present invention are:

[0020] The present invention performs lens cutting on the electronic evidence information to be identified to obtain multiple lens information, and then determines whether each lens information has been tampered with, so as to ensure that the subsequent evidence-based analysis and judgment are established on a reliable foundation, thereby improving the credibility of the entire electronic evidence processing flow. When the electronic evidence information to be identified has not been tampered with, the electronic evidence information to be identified is screened to obtain a screening result, and finally the illegal behavior of the target object is identified based on the screening result, thereby improving the efficiency of electronic evidence in identifying illegal behavior.

[0021] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by implementing the embodiments of the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 The figure is a flow chart of the electronic evidence identification method in the present invention.

[0024] Figure 2 It is a schematic diagram of the topological structure of the electronic evidence identification system in the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0027] Embodiment 1:

[0028] This embodiment provides an electronic evidence identification method based on a neural network. It can be understood that a scenario can be set up in this embodiment, for example: processing the video stream data collected by a surveillance camera to determine whether there is a scenario in which illegal behavior exists in the video stream data.

[0029] See also Figure 1 , the figure shows that the method includes step S1, step S2, step S3, step S4 and step S5.

[0030] Step S1, obtaining electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information;

[0031] Step S2, performing shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two consecutive image frames;

[0032] The video stream is composed of frames, shots, and scenes, which are continuously stacked and spliced ​​in chronological order. The video stream information is cut into shots to obtain shot information, and continuous image frames are analyzed as a whole, which provides a more detailed and accurate basic unit for subsequent judgment and analysis, so that the processing of video content is not based on scattered images, but on coherent shots, which can better reflect the actual content of the video and scene changes.

[0033] Step S3, judging whether the electronic evidence information to be identified has been tampered with according to the lens information, and obtaining a first judgment result;

[0034] In the application of electronic evidence, the authenticity and integrity of the evidence are of vital importance. This step uses the lens information to determine whether the electronic evidence information to be identified has been tampered with. It can promptly detect human modifications to the video evidence and effectively exclude tampered video evidence, ensuring that subsequent evidence-based analysis and judgment are based on a reliable foundation, thereby improving the credibility of the entire electronic evidence processing process.

[0035] Step S4: when the first judgment result is that the electronic evidence information to be identified has not been tampered with, the electronic evidence information to be identified is screened to obtain a screening result, wherein the screening result includes an image frame in which the target action exists;

[0036] Further, the step S4 includes step S41, step S42 and step S43:

[0037] Step S41, sending the image frames in the electronic evidence information to be identified to a human posture estimation model to obtain skeleton position information corresponding to each image frame;

[0038] Step S42, extracting features from the skeleton position information corresponding to each image frame to obtain feature sequence information;

[0039] Step S43: sending the feature sequence information to the spatiotemporal feature long short-term memory network to obtain a screening result.

[0040] In this step, the temporal dependency and spatial dependency between human joints can be modeled based on the spatiotemporal feature long short-term memory network to complete the detection of the target action, so that the key images related to the target action can be quickly located from a large amount of video information, avoiding indiscriminate analysis of the entire video, greatly reducing the amount of data processing, improving work efficiency, and highlighting the key analysis content.

[0041] Step S5: Identify the illegal behavior of the target object according to the screening results to obtain an identification result, so that the determination of the illegal behavior is more accurate and objective, which helps to improve the fairness and accuracy of law enforcement.

[0042] Embodiment 2:

[0043] The difference between this embodiment and embodiment 1 is that step S3, judging whether the electronic evidence information to be identified has been tampered with according to the lens information, and obtaining a first judgment result, includes steps S31 to S33:

[0044] Step S31, determining the key frame included in each shot information according to the shot information, which specifically includes the following steps:

[0045] Step S311, acquiring two adjacent image frames;

[0046] Step S312, calculating the mean, variance and covariance in the local window of two adjacent image frames to obtain a calculation result;

[0047] Step S313, determining the similarity between two adjacent image frames according to the calculation result to obtain similarity information;

[0048] In this step, the specific calculation formula of similarity is:

[0049]

[0050] In the above formula, S represents similarity; α 1 and α 2 Respectively represent the mean values ​​in the local windows of two adjacent image frames; and Respectively represent the variance in the local window of two adjacent image frames; σ 12 represents the covariance of the local window of two adjacent image frames, c 1 and c 2 Represents a constant.

[0051] Step S314, determine the key frame according to the similarity information of each two adjacent image frames. In this step, by calculating the similarity between each two adjacent image frames, the frame sequence can be used as the horizontal axis and the similarity can be used as the vertical axis to draw a similarity curve, and the similarity curve is smoothed and denoised. The processed similarity curve can measure whether there is a large difference between adjacent frames to extract the key frame.

[0052] Step S32: preprocess the key frame to obtain a preprocessed key frame, wherein the preprocessing includes enhancing the key frame to improve image quality and enrich information.

[0053] Step S33: judging whether the electronic evidence information to be identified has been tampered with according to the preprocessed key frame, and obtaining a first judgment result.

[0054] When judging whether electronic evidence has been tampered with, operations such as editing and splicing will leave more obvious traces on the key frames. By capturing the key frames included in each shot information and then preprocessing the key frames, the characteristics of the key frames are further optimized, making subsequent tampering judgments based on the preprocessed key frames more accurate, and improving the ability to identify tampering features, thereby more accurately judging the authenticity of electronic evidence.

[0055] Specifically, the step S33 includes steps S331 to S335:

[0056] Step S331, extracting feature information corresponding to the preprocessed key frame in the spatial domain to obtain first feature information;

[0057] In this step, the preprocessed key frame can be sent to the multi-scale fusion network for processing multiple times to obtain the first feature information, wherein the multi-scale fusion network processing includes firstly increasing the dimension through a 1×1 convolution layer, and then reducing the amount of calculation through a 3×3 DW convolution, and then using dilated convolutions with different dilation factors to capture multi-scale tampering detail information and splice it on the channel, and then learning the importance of each scale feature channel through global average pooling and two fully connected layers to obtain the local discontinuity of the preprocessed key frame, and then reducing the dimension through a 1×1 convolution, and then preventing overfitting through a 1×1 dimensionality reduction PW convolution and a random inactivation layer to obtain forged features, and finally adding it to itself and forming the final multi-scale fusion feature, i.e., the first feature information, through a channel-space attention module.

[0058] Step S332, extract the feature information corresponding to the preprocessed key frame in the frequency domain to obtain the second feature information. For example, in this step, the frequency domain image can be obtained through the high-frequency noise flow, and then the frequency domain image can be sent to the multi-scale fusion network multiple times to obtain the high-frequency noise information, that is, the second feature information.

[0059] Step S333: concatenate the first feature information with the second feature information to obtain concatenated feature information;

[0060] Step S334, sending the spliced ​​feature information to the attention network to obtain weighted feature information;

[0061] In this step, the channel attention mechanism can be used to calculate the channel dimension information, and then the spatial attention mechanism can be used to calculate the spatial dimension information and perform residual connection with the spliced ​​feature information to obtain the interactive features. Finally, the original feature map and the weighted feature map are added to obtain the influence of the first feature information and the second feature information on the tampered information in their respective channels.

[0062] Step S335: determine whether the electronic evidence information to be identified has been tampered with based on the weighted feature information to obtain a first judgment result.

[0063] When the attention mechanism is used to fuse the first feature information in the spatial domain and the second feature information in the frequency domain, the feature streams of two different modalities can interact with each other and promote learning, rather than keeping them independent of each other, thereby capturing the inconsistency between the tampered area and the real area, and accurately judging whether the electronic evidence information to be identified has been tampered with.

[0064] Embodiment 3:

[0065] The difference between this embodiment and embodiment 1 or 2 is that step S5, identifying the illegal behavior of the target object according to the screening result to obtain the identification result, includes steps S51 to S54:

[0066] Step S51, sending the screening result to a spatial feature extraction network to obtain third feature information;

[0067] In surveillance scenes with a large field of view, small human targets, crowds and occlusions make it difficult for traditional methods to grasp the spatial relationship between human objects. The spatial feature extraction network can focus on the spatial structure of the image and conduct in-depth analysis and extraction of the spatial position, shape and relative position relationship of the target objects in the screening results. This helps to capture the potential relationship between human objects and avoid misjudgment of the target objects due to the loss or confusion of local details, thereby providing a more accurate spatial information basis for subsequent illegal behavior identification.

[0068] Step S52: sending the electronic evidence information to be identified to the spatiotemporal feature extraction network to obtain fourth feature information;

[0069] In surveillance videos, the behavior of the target object changes over time, and there is a dynamic interaction between the object and the environment, and between the environment and the environment. The spatiotemporal feature extraction network can fully consider the temporal dimension of the video, not only extracting the spatial features in each frame, but also capturing the target object's motion trajectory, behavioral changes, and interaction information with the surrounding environment at different times. This helps to better understand the behavior patterns of the target object and their relationship with the surrounding environment in complex surveillance scenarios, overcoming the problem that traditional methods easily separate the potential relationship between the object and the environment, and between the environment and the environment.

[0070] Step S53, sending the third feature information and the fourth feature information to a feature fusion attention network to obtain fused feature information;

[0071] The feature fusion attention network can deeply fuse feature information from different channels, and use the attention mechanism to automatically learn the importance of different features. In complex monitoring scenarios, there is a large amount of information, some of which is crucial to the identification of illegal behavior, while some may be interference information. The attention mechanism can focus on key features, give higher weights to these important features, and suppress irrelevant or minor features. In this way, the most representative and discriminative information can be extracted from the fused feature information, further improving the accuracy of identifying illegal behaviors of target objects in complex scenarios.

[0072] Step S54: Identify the illegal behavior of the target object based on the fused feature information to obtain an identification result.

[0073] Specifically, the step S54 includes steps S541 to S545:

[0074] Step S541, determining whether there is a target object corresponding to the target action in the screening results;

[0075] Step S542, segmenting the face image of the target object to obtain face image information;

[0076] Step S543, determining whether the facial image information is blocked, and obtaining a second determination result;

[0077] In most scenarios, the face may be obstructed by objects such as sunglasses, scarves, and masks, which may make it impossible to accurately identify the target object's emotions, thereby reducing the accuracy of illegal behavior identification. Therefore, when identifying illegal behavior, it is necessary to first determine whether the target object's facial image information is obstructed.

[0078] Step S544, identifying the emotion of the target object according to the second judgment result to obtain an emotion recognition result;

[0079] In this step, when the second judgment result is that there is occlusion in the facial image of the target object, the occluded part in the facial image information is determined to obtain the occluded image information; the occluded image information and the facial image information are sent to a partial convolutional generative adversarial network to obtain a de-occluded facial image; the emotion of the target object is identified based on the de-occluded facial image to obtain an emotion recognition result, and the emotion of the target object is accurately determined by repairing the occluded facial image, thereby improving the applicability of the present invention and making it applicable to most scenarios.

[0080] Step S545: Identify the illegal behavior of the target object based on the emotion recognition result and the fused feature information.

[0081] Identifying illegal behavior based only on the target action of a single-modal target object may result in misidentification. Therefore, in this embodiment, identifying the illegal behavior of the target object in combination with the emotion recognition result of the target object can effectively reduce the misidentification rate of illegal behavior.

[0082] Embodiment 4:

[0083] like Figure 2 As shown, this embodiment provides an electronic evidence recognition system based on a neural network, the system includes an acquisition module 901, a first processing module 902, a judgment module 903, a second processing module 904 and a third processing module 905, which specifically include:

[0084] An acquisition module 901 is used to acquire electronic evidence information to be identified, where the electronic evidence information to be identified includes video stream information;

[0085] A first processing module 902 is used to perform shot cutting on the video stream information to obtain at least one shot information, where the shot information includes at least two consecutive image frames;

[0086] A judgment module 903 is used to judge whether the electronic evidence information to be identified has been tampered with according to the lens information, and obtain a first judgment result;

[0087] A second processing module 904 is configured to, when the first judgment result is that the electronic evidence information to be identified has not been tampered with, filter the electronic evidence information to be identified to obtain a filtering result, wherein the filtering result includes an image frame in which a target action exists;

[0088] The third processing module 905 is used to identify the illegal behavior of the target object according to the screening result to obtain an identification result.

[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0090] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for identifying electronic evidence, characterized in that: include: Acquiring electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information; Performing shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two continuous image frames; Determining whether the electronic evidence information to be identified has been tampered with according to the lens information, and obtaining a first determination result; When the first judgment result is that the electronic evidence information to be identified has not been tampered with, screening the electronic evidence information to be identified to obtain a screening result, wherein the screening result includes an image frame in which the target action exists; The illegal behavior of the target object is identified according to the screening result to obtain an identification result.

2. The electronic evidence identification method according to claim 1, characterized in that: Determining whether the electronic evidence information to be identified has been tampered with according to the lens information to obtain a first determination result includes: Determine the key frame included in each shot information according to the shot information; Preprocessing the key frame to obtain a preprocessed key frame; It is determined whether the electronic evidence information to be identified has been tampered with according to the preprocessed key frame to obtain a first determination result.

3. The electronic evidence identification method according to claim 2, characterized in that: Determining whether the electronic evidence information to be identified has been tampered with according to the preprocessed key frame to obtain a first determination result includes: Extracting feature information corresponding to the preprocessed key frame in the spatial domain to obtain first feature information; Extracting feature information corresponding to the preprocessed key frame in the frequency domain to obtain second feature information; Concatenate the first characteristic information with the second characteristic information to obtain concatenated characteristic information; Sending the concatenated feature information to an attention network to obtain weighted feature information; It is determined whether the electronic evidence information to be identified has been tampered with according to the weighted characteristic information to obtain a first determination result.

4. The electronic evidence identification method according to claim 2, characterized in that: Determining a key frame included in each shot information according to the shot information includes: Acquire two adjacent image frames; Calculate the mean, variance and covariance in the local window of two adjacent image frames to obtain the calculation result; Determine the similarity between two adjacent image frames according to the calculation result to obtain similarity information; The key frame is determined according to the similarity information between each two adjacent image frames.

5. The electronic evidence identification method according to claim 1, characterized in that: The illegal behavior of the target object is identified according to the screening result to obtain an identification result, including: Sending the screening result to a spatial feature extraction network to obtain third feature information; Sending the electronic evidence information to be identified to the spatiotemporal feature extraction network to obtain fourth feature information; Sending the third feature information and the fourth feature information to a feature fusion attention network to obtain fused feature information; The illegal behavior of the target object is identified according to the fused feature information to obtain an identification result.

6. The electronic evidence identification method according to claim 5, characterized in that: The illegal behavior of the target object is identified according to the fused feature information to obtain an identification result, including: Determine whether there is a target object corresponding to the target action in the screening results; Segmenting the face image of the target object to obtain face image information; Determine whether the facial image information is blocked, and obtain a second determination result; Recognize the emotion of the target object according to the second judgment result to obtain an emotion recognition result; The illegal behavior of the target object is identified according to the emotion recognition result and the fused feature information.

7. The electronic evidence identification method according to claim 6, characterized in that: The emotion of the target object is identified according to the second judgment result to obtain an emotion recognition result, including: When the second judgment result is that the face image of the target object is blocked, determining the blocked part in the face image information to obtain blocked image information; Sending the occluded image information and the face image information to a partial convolutional generative adversarial network to obtain a deoccluded face image; The emotion of the target object is recognized according to the de-occluded face image to obtain an emotion recognition result.

8. The electronic evidence identification method according to claim 4, characterized in that: The specific calculation method of similarity is: In the above formula, S represents similarity; α1 and α2 represent the mean values ​​in the local windows of two adjacent image frames respectively; and Respectively represent the variance in the local window of two adjacent image frames; σ 12 represents the covariance within the local window of two adjacent image frames, and c1 and c2 represent constants.

9. The electronic evidence identification method according to claim 1, characterized in that: The electronic evidence information to be identified is screened to obtain screening results, including: Sending the image frames in the electronic evidence information to be identified to a human posture estimation model to obtain skeleton position information corresponding to each image frame; Extract features from the skeleton position information corresponding to each image frame to obtain feature sequence information; The feature sequence information is sent to the spatiotemporal feature long short-term memory network to obtain a screening result.

10. An electronic evidence recognition system based on a neural network, characterized in that: include: An acquisition module, used for acquiring electronic evidence information to be identified, wherein the electronic evidence information to be identified includes video stream information; A first processing module, configured to perform shot cutting on the video stream information to obtain at least one shot information, wherein the shot information includes at least two continuous image frames; A judgment module, used for judging whether the electronic evidence information to be identified has been tampered with according to the lens information, and obtaining a first judgment result; A second processing module is configured to, when the first judgment result is that the electronic evidence information to be identified has not been tampered with, screen the electronic evidence information to be identified to obtain a screening result, wherein the screening result includes an image frame in which a target action exists; The third processing module is used to identify the illegal behavior of the target object according to the screening result to obtain an identification result.

Citation Information

Patent Citations

  • Motor vehicle violation shooting system and method

    CN106408952A

  • Tampered video identification method and device, electronic equipment and storage medium

    CN118799772A