Image display method, system, and electronic device
By displaying images marked with tampered areas, the problem of artificial intelligence forgery or tampering with image recognition is solved, reminding and accurate recognition of image content tampering is achieved, and the accuracy of image recognition and user operation flexibility is improved.
Patent Information
- Application Number
- PCT/CN2024/139174
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art is difficult to effectively identify and process images that have been forged or tampered with by artificial intelligence, resulting in users being unable to accurately perform image recognition and subsequent operations.
By displaying images marked with tampered areas, the user can clearly identify whether the image has been tampered with and the tampered areas, thereby reminding the user to perform tampering inspections of image content.
It realizes reminders of image content tampering, helps users to more accurately identify information in the image, and improves the accuracy of image recognition and user operation flexibility.
Smart Images

Figure CN2024139174_19062025_PF_FP_ABST
Abstract
Description
Image display method, system and electronic device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 13, 2023, with application number 202311714687.0 and invention name “A method, system and electronic device for image recognition”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of intelligent recognition technology, and in particular to an image display method, system and electronic equipment. Background Art
[0003] With the continuous advancement of AI (Artificial Intelligence) forgery technology, many images have been forged or tampered with by AI, resulting in inconsistencies between the image content and the actual content. Since users cannot determine whether the image content has been tampered with, they may be unable to perform subsequent image processing or recognition operations. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide an image display method, system, and electronic device to alert users of tampering with image content, facilitating subsequent operations on the image. The specific technical solution is as follows:
[0005] The present invention provides an image display method, which includes:
[0006] The first image is displayed with the tampered area marked.
[0007] The present application also provides an image recognition system, comprising:
[0008] The business processing unit is configured to display the first image marked with the tampered area.
[0009] An embodiment of the present application further provides an electronic device, including:
[0010] Memory for storing computer programs;
[0011] The processor is configured to implement any of the above-mentioned image display methods when executing a program stored in the memory.
[0012] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned image display methods is implemented.
[0013] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the above-described image display methods.
[0014] Beneficial effects of the embodiments of the present application:
[0015] An image display method, system, and electronic device provided in the embodiments of the present application can, by displaying a first image with annotated areas, enable a user to more clearly distinguish whether the first image has been tampered with and the tampered areas in the first image, thereby reminding the user of tampering with the image content, thereby facilitating subsequent operations on multimedia data by the user, and further facilitating the user to more accurately obtain more accurate information from the multimedia data.
[0016] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0018] FIG1 is a schematic diagram of a flow chart of an image display method provided by the present application;
[0019] FIG2a is a schematic diagram of a display interface provided by the present application when a first image is tampered with;
[0020] FIG2b is a schematic diagram of a display interface provided by the present application when the first image has not been tampered with;
[0021] FIG2c is another schematic diagram of the display interface when the first image provided by the present application is tampered with;
[0022] FIG2d is a schematic diagram of a display interface provided by the present application to indicate the progress of tamper detection;
[0023] FIG3 is a flow chart of an image recognition method provided by the present application based on the image display method;
[0024] FIG4a is a schematic diagram of a first image with a non-tampering area marked thereon provided by the present application;
[0025] FIG4b is a schematic diagram of a first image for determining an area to be identified provided by the present application;
[0026] FIG4c is a schematic diagram of the display process of the attribute selection interface provided by this application;
[0027] FIG5 is another flowchart of an image recognition method provided by the present application based on the image display method;
[0028] FIG6 is a schematic diagram of a spliced display of a tampered first image provided by the present application;
[0029] FIG7 is a schematic diagram of another flow chart of an image recognition method provided by the present application based on the image display method;
[0030] FIG8 is a schematic diagram of the structure of an image recognition system provided by this application;
[0031] FIG9 is a schematic diagram of another flow chart of an image recognition method provided by the present application based on the image display method;
[0032] FIG10 is a schematic structural diagram of an image display device provided by the present application;
[0033] FIG11 is a schematic structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION
[0034] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described below with reference to the accompanying drawings and examples. It should be understood that the described examples are only some examples of the present invention, not all examples. All other examples derived by persons of ordinary skill in the art based on the examples of the present invention fall within the scope of protection of the present invention.
[0035] In order to remind users of image content tampering and facilitate subsequent operations on the image, the present invention provides an image display method, system, and electronic device. The following first describes the image display method provided by the present invention.
[0036] The image display method provided in an embodiment of the present application includes: displaying a first image marked with a tampered area.
[0037] By applying the embodiments of the present application, by displaying the first image with the tampered area marked, the user can more clearly distinguish whether the first image has been tampered with and the tampered area in the first image, thereby reminding the user of the tampering of the image content, thereby facilitating the user's subsequent operations on the multimedia data, and further facilitating the user to more accurately obtain more accurate information from the multimedia data.
[0038] Specifically, as shown in FIG1 , the image display method provided by the present application may include the following steps:
[0039] S101 , in response to a play instruction inputted to a display interface, obtaining an image included in multimedia data indicated by the play instruction as a first image, and obtaining a tampered area obtained by pre-tampering detection on the first image.
[0040] S102: Displaying a first image with a tampered area marked in a multimedia data playback window of a display interface.
[0041] With this embodiment, in response to a play instruction input to a display interface, an image included in the multimedia data indicated by the play instruction is obtained as a first image, and a tampered region obtained by pre-tampering detection of the first image is obtained; the first image with the tampered region marked is displayed in the multimedia data playback window of the display interface. By displaying the first image with the tampered region marked, the user can more clearly distinguish whether the first image has been tampered with and the tampered region in the first image, thereby reminding the user of image content tampering, thereby facilitating the user's subsequent operations on the multimedia data and enabling the user to more accurately obtain more accurate information from the multimedia data.
[0042] The following will provide exemplary descriptions of the aforementioned S101-S102:
[0043] In S101, the display interface may be as shown in FIG2a, and the user may input a play instruction by clicking on the multimedia data shown in the display interface; the user may also input a play instruction by clicking on the multimedia data to be played and setting the playback method through the display interface.
[0044] The multimedia data may refer to an image, in which case the first image is the multimedia data itself. The multimedia data may also refer to a video, in which case the first image is each video frame contained in the multimedia data.
[0045] Specifically, if the multimedia data is a video, then multiple video frames in the multimedia data can be used as the first image to perform tampering detection on the first image, that is, tampering detection is performed on multiple video frames in the multimedia data. The multiple video frames in this example can refer to all video frames in the video, or can refer to multiple video frames in the video that contain information of interest to the user. If the multimedia data is a video, then only one video frame in the multimedia data can be used as the first image to perform tampering detection on the first image, that is, tampering detection is performed on only one video frame in the multimedia data, which is a video frame that contains information of interest to the user. If the tampering detection result of the first image is tampered when tampering detection is performed in advance, the tampered area in the tampered first image can be obtained. The following will provide an exemplary explanation of how to perform tampering detection on the first image, which will not be repeated here.
[0046] In S102, referring to FIG2a, the display interface includes a multimedia data playback window. The multimedia data playback window displays the image included in the multimedia data indicated by the playback instruction, that is, the multimedia data playback window displays a first image. If the tampering detection result of the first image is tampered, the displayed first image will be marked with a tampered area. For example, as shown in FIG2a, the multimedia data playback window displays that the first image has been tampered with, and the tampered area is the area surrounded by a rectangular frame, that is, the tampered area is a person's face. The rectangular frame is a tampering mark used to mark the tampered area in the first image. In other possible embodiments, the tampering mark can also be a circular frame, etc.
[0047] It is understandable that the tampering detection result of the first image can be either tampered or not tampered. To facilitate the user to distinguish between tampered and untampered first images, in one possible embodiment, if the multimedia data is a video and the tampering detection result of the first image is tampered, as shown in FIG2a, while displaying the tampered area in the first image, the words "Video has been tampered" can also be displayed in the lower right corner of the multimedia data playback window. If the tampering detection result of the first image is not tampered, as shown in FIG2b, the words "Video has not been tampered" can be displayed in the lower right corner of the multimedia data playback window.
[0048] In a possible embodiment, in response to a tampering detection result that the first image displayed in the multimedia data playback window is tampered with, a tampering annotation display control may be displayed in the display interface.
[0049] In response to a tampering annotation hiding operation on a tampering annotation display control, hiding the annotation of the tampered area in the first image; or, the aforementioned S102 includes: in response to a tampering annotation display operation on a tampering annotation display control, displaying the annotation of the tampered area in the first image.
[0050] If the tampering detection result of the first image displayed in the multimedia data playback window is tampered, a tampering mark display control can be displayed in the display interface, so that the user can control whether to display the mark of the tampered area according to the tampering mark display control. If the tampering detection result of the first image displayed in the multimedia data playback window is not tampered, there is no tampered area in the first image, and there is no mark of the tampered area in the first image, so there is no need to display the tampering mark display control.
[0051] For example, the style of the tampering mark display control can be as shown in the logo after the prompt "The video has been tampered" in Figure 2a. The tampering mark display control in Figure 2a is in the open state, that is, the multimedia data playback window will display the mark of the tampered area in the first image, that is, the tampered area mark of the rectangular frame shown in Figure 2a.
[0052] When the user clicks on the tampering annotation display control, so that the tampering annotation display control is in the open state, the style of the tampering annotation display control can be as shown in Figure 2a, and the multimedia data playback window will display the annotation of the tampered area in the first image, that is, display the tampering annotation, so that the user can distinguish between the tampered area and the non-tampered area in the first image.
[0053] When the user clicks on the tampering annotation display control, the tampering annotation display control is in the closed state. At this time, the style of the tampering annotation display control can be as shown in the style of hiding the tampering annotation in Figure 2c. Then, the multimedia data playback window will not display the annotation of the tampered area in the first image, that is, the annotation of the tampered area in the first image is hidden.
[0054] In this example, the operation of the user clicking on the tampering annotation display control so that the tampering annotation display control is in an on state can be regarded as a tampering annotation display operation, and the operation of the user clicking on the tampering annotation display control so that the tampering annotation display control is in an off state can be regarded as a tampering annotation hiding operation. In a possible embodiment, the tampering annotation display control can be in an on state by default to help the user distinguish between the tampered area and the non-tampered area in the first image. If the user needs to hide the annotation of the tampered area in the first image, the user can click on the tampering annotation display control so that the tampering annotation display control is in an off state. In response to the user's tampering annotation hiding operation, the multimedia data playback window of the display interface hides the annotation of the tampered area in the first image.
[0055] By selecting this embodiment, when the first image is tampered with, the user can control whether the multimedia data playback window displays the annotation of the tampered area in the first image through the tampering annotation display control displayed on the display interface, so that the user can control the display of the annotation of the tampered area in the first image according to his own needs, thereby improving the user's operating flexibility on the display interface.
[0056] In a possible embodiment, the display interface also includes: a multimedia data list, the multimedia data list includes identifiers corresponding to multiple multimedia data; the method also includes: if the multimedia data includes a tampered image, setting a tampering label for the identifier corresponding to the multimedia data in the multimedia data list.
[0057] In this embodiment, if the tampering detection result of an image included in a certain multimedia data is tampered, the multimedia data is considered to be tampered, and the label of the multimedia data can be set as a tampering label, which will be displayed after the corresponding identifier of the multimedia data in the multimedia data list.
[0058] The tampering tag is a tag used to indicate whether the multimedia data has been tampered with. Specifically, the tampering tag can be displayed in various forms, such as text, symbols, etc. For example, if the tampering tag is in text form, the tampered multimedia data can be labeled as "tampered". If the tampering tag is in symbol form, the tampered multimedia data can be labeled as an "×" symbol.
[0059] It is understandable that the multimedia data in the multimedia data list may be multimedia data that has been pre-tampered detected or multimedia data that has not yet been tampered detected. The tampering detection result of multimedia data that has been pre-tampered detected may be tampered or not tampered. In one possible implementation, no distinction may be made between multimedia data that has not been tampered with and multimedia data that has not yet been tampered detected. For example, using the tampering label in text form as an example, the label of tampered multimedia data may be set as "tampered", while multimedia data that has not been tampered with and multimedia data that has not yet been tampered detected may not be set with a label.
[0060] In another possible implementation, in order to facilitate users to better determine whether multimedia data has been tampered with, it is also possible to distinguish between multimedia data that has not been tampered with and multimedia data that has not been detected for tampering. For example, taking the tampering label in text form as an example, the label of the tampered multimedia data can be set to "tampered", and the label of the tampered multimedia data can be set to "not tampered", and the multimedia data that has not been detected for tampering is not set with a label. Taking the tampering label in symbol form as an example, the label of the tampered multimedia data can be set to the symbol "×", and the label of the tampered multimedia data can be set to the symbol "√", and the multimedia data that has not been detected for tampering is not set with a label.
[0061] The multimedia data list displays identifiers corresponding to multiple multimedia data, wherein the multiple multimedia data may refer to multimedia data that the user has played, or may refer to multimedia data that the user has not played yet. For example, referring to the display interface shown in FIG2a, the "playback history" in FIG2a is the multimedia data list, and the identifiers of video 1 to video 8 are the identifiers corresponding to the 8 multimedia data respectively. In this example, the multimedia data displayed in the multimedia data list is the multimedia data that the user has played. Video 2, video 3, and video 4 with a tampering label of "tampered" are tampered multimedia data; video 5 and video 6 with a tampering label of "not tampered" are not tampered multimedia data; video 1, video 7, and video 8 without a tampering label are multimedia data that have not been detected for tampering.
[0062] By selecting this embodiment, the tampering label in the list can more clearly remind the user whether the multimedia data has been tampered with, thereby facilitating the user to perform subsequent operations on the multimedia data.
[0063] In the case where the multimedia data is a video clip and the first image is each video frame in the video clip, in order to facilitate the user to more clearly distinguish the tampered video frames and the non-tampered video frames in the video clip, in a possible embodiment, tampering detection can be performed on each video frame according to the order of each video frame in the video clip, and the video frame with the tampering detection result being tampered with in each video frame is determined as the tampered video frame; the time area corresponding to the tampered video frame in the time progress bar of the video clip is determined, and a tampering mark is set in the time area.
[0064] According to the order of each video frame in the video clip, tampering detection is performed on each video frame in the video clip one by one to obtain a tampering detection result of each video frame. If the tampering detection result of a video frame is tampered, the video frame is used as a tampered video frame.
[0065] According to the time set corresponding to each tampered video frame in the video clip, the time period corresponding to all tampered video frames in the video clip can be determined, and then the time area corresponding to all tampered video frames in the time progress bar of the video clip can be determined. A tampering mark can be set at the determined time area so that the user can distinguish the tampered part and the non-tampered part in the video clip through the tampering mark on the progress bar.
[0066] Tamper indicators can be displayed in various forms, such as color and shape. For example, if the tamper indicator is in the form of color, the tamper indicator should be a color different from the color of the progress bar itself. For example, if the progress bar is black, the tamper indicator should be orange. If the tamper indicator is in the form of a shape, the tamper indicator should be a shape different from the shape of the progress bar itself. For example, if the progress bar is an unfilled rectangle, the tamper indicator can be a grid-filled rectangle.
[0067] For example, the progress bar of the multimedia data playback window when playing a video can be shown in Figure 2a. Referring to Figure 2a, the shape of the progress bar is a rectangle filled with vertical lines, the tampering mark is a rectangle filled with diagonal lines, and the circular mark is the position of the progress bar corresponding to the video frame currently played in the multimedia data playback window. In this example, the shape of the progress bar is a rectangle filled with vertical lines, and the time area corresponding to the tampered video frame is a rectangle filled with diagonal lines. As the multimedia data playback window progresses in playing the video, the circular mark moves on the progress bar, and the shape of the area corresponding to the video frame that has been played in the progress bar will become an unfilled rectangle. However, in order to distinguish the tampered part from the untampered part in the video clip, the area corresponding to the video frame that has been played and tampered in Figure 2a in the progress bar is still displayed in the shape of the tampering mark, that is, it is displayed as a rectangle filled with diagonal lines.
[0068] If the video clip played in the multimedia data playback window has not been pre-detected for tampering, the progress bar will be a rectangle filled with vertical lines. As the multimedia data playback window progresses in playing the video, the circular mark moves on the progress bar, and the shape of the area corresponding to the video frames that have been played in the progress bar will become an unfilled rectangle.
[0069] If only some of the video frames in the video clip played in the multimedia data playback window have been pre-detected for tampering, then the shape of the area corresponding to the tampered video frames that have been detected for tampering and whose tampering detection results are tampered is a rectangle filled with diagonal lines in the progress bar. As the multimedia data playback window progresses in playing the video, the circular mark moves on the progress bar, and the shape of the area corresponding to the video frames that have been played and whose tampering detection results are not tampered with in the progress bar will all change to an unfilled rectangle, while the shape of the area corresponding to the video frames that have been played and whose tampering detection results are tampered with in the progress bar will still be a rectangle filled with diagonal lines, and the shape of the area corresponding to the video frames that have been played but have not been detected for tampering in the progress bar will be an unfilled rectangle. If a video frame that has been played but has not been detected for tampering is detected for tampering and the tampering detection result is tampered, the shape of the area corresponding to the video frame in the progress bar will change from an unfilled rectangle to a rectangle filled with diagonal lines.
[0070] By selecting this embodiment, tampering detection can be performed on each video frame in the order of each video frame in the video clip, and the video frame with the tampering detection result being tampered in each video frame can be determined as the tampered video frame; the time area corresponding to the tampered video frame in the time progress bar of the video clip is determined, and a tampering mark is set in the time area, so that the user can more clearly judge the tampered video frame and the non-tampered video frame in the video clip through the tampering mark on the progress bar.
[0071] In the case where the multimedia data is a video clip and the first image is each video frame in the video clip, if the multiple video clips have not been pre-tampered with, the video clip can be tampered with when the video clip is played in the multimedia data playback window. During the tampering detection process, in order to facilitate the user to view the tampering detection progress of the currently playing video clip, in one possible embodiment, tampering detection is performed on each video frame in the order of each video frame in the video clip, and the tampering detection progress of the video clip is displayed in the display interface; the tampering detection progress is the ratio of the number of video frames in the video clip that have been tampered with to the number of all video frames included in the video clip.
[0072] The tampering detection progress is the ratio of the number of video frames in a video clip that have been tampered with to the total number of video frames in the video clip. Since there is a corresponding relationship between the number of video frames and the playback duration, the tampering detection progress can also be considered as the ratio of the playback duration of the video frames in a video clip that have been tampered with to the total duration of the video clip.
[0073] Since tampering detection is performed on each video frame in the video clip one by one in the order of each video frame in the video clip, if the playback time of the video frame that has been tampered with is 5 minutes and the total length of the video clip is 10 minutes, the tampering detection progress is 5 / 10×100%=50%. Therefore, the user can determine from the tampering detection progress that the first 5 minutes of the currently playing video clip have been tampered with detection, and the last 5 minutes have not been tampered with detection.
[0074] For example, as shown in FIG2 d , the words “0% in tampering detection” shown in the lower right corner of the multimedia data playback window are prompts of the tampering detection progress of the currently playing video clip.
[0075] When tampering detection is performed on the currently played video segment, the display form of the scan lines in FIG. 2 d may also be used to indicate that tampering detection is currently being performed on the video frame.
[0076] By selecting this embodiment, the user can be reminded of the tampering detection progress of the currently playing video clip by displaying the tampering detection progress of the video clip in the display interface, so that the user can distinguish between video frames that have been tampered with and video frames that have not been tampered with, thereby making it easier for the user to obtain information on whether the video frame has been tampered with.
[0077] When viewing an image displayed on a display interface, users can obtain related or similar image information by identifying the content within the image. This method of obtaining image information is hereinafter referred to as image recognition. However, because the content within the image has been tampered with, directly recognizing the image may result in information related to the tampered image content, i.e., information irrelevant to the original image content, thereby reducing the accuracy of image recognition. Furthermore, the image information obtained through direct image recognition is relatively broad, which may prevent users from obtaining the image information of interest, resulting in reduced image recognition accuracy.
[0078] Based on this, in order to make the image recognition results more accurate, based on the above-mentioned image display method, this application also provides an image recognition method, as shown in FIG3 , the method includes:
[0079] S301: Determine a non-tampered area in a first image.
[0080] S302: Display the first image with the non-tampering area marked.
[0081] S303 : In response to the region selection operation on the first image, determine, in the first image, a region to be identified that is selected by the region selection operation.
[0082] S304: Determine candidate search attributes possessed by the object existing in the area to be identified from the preset search attributes.
[0083] S305: Display an attribute selection interface including candidate search attributes.
[0084] S306 : In response to the attribute selection operation on the attribute selection interface, identifying the candidate search attribute selected by the attribute selection operation as the target search attribute.
[0085] S307: Acquire the attribute value of the target retrieval attribute of the area to be identified as the target attribute value.
[0086] S308 : Searching for an object whose attribute value of the target retrieval attribute matches the target attribute value in a database that pre-stores various objects and attribute values of preset retrieval attributes of various objects, as a recognition result.
[0087] By selecting this embodiment, a non-tampered area in a first image that has not been tampered with can be determined, and the first image marked with the non-tampered area can be displayed. In response to an area selection operation on the first image, an area to be identified selected by the area selection operation can be determined in the first image; an alternative search attribute possessed by an object in the area to be identified can be determined in the preset search attributes; an attribute selection interface containing the alternative search attributes can be displayed; in response to an attribute selection operation on the attribute selection interface, the alternative search attribute selected by the attribute selection operation can be identified as a target search attribute, and an attribute value of the target search attribute in the area to be identified can be obtained as a target attribute value. In a database pre-stored with attribute values of various objects and preset search attributes of various objects, an object whose attribute value of the target search attribute matches the target attribute value can be searched for as a recognition result. When determining the target search attribute, the alternative search attribute possessed by the object in the area to be identified can be determined in the preset search attributes. Since the alternative search attribute can be regarded as a preset search attribute that a user may select, the alternative search attribute can be recommended as a preset search attribute for the user to select, thereby helping the user quickly determine the target search attribute possessed by the content of interest to the user from multiple preset search attributes. Since the database is established based on the attribute values of the preset retrieval attributes of each object, and the target retrieval attribute for the area to be identified is determined from the preset retrieval attributes, and the target attribute value is the attribute value of the target retrieval attribute of the area to be identified, it is possible to screen the objects in the database by matching the attribute values of the target retrieval attributes of each object in the database with the target attribute value, and use the objects whose attribute values of the target retrieval attributes in the database match the target attribute value as the recognition results, thereby more accurately obtaining relevant information about the content that the user is interested in. Therefore, the embodiments of the present application can determine the characteristics of the content that the user is interested in by determining the target retrieval attribute, and more accurately determine relevant information about the content that the user is interested in by matching the attribute values of the target retrieval attributes of each object in the database with the target attribute value, thereby improving the accuracy of image recognition.
[0088] The above S301-S308 will be described below respectively, wherein:
[0089] In S301, assume that a first image contains objects A, B, and C, where some features of object A have been tampered with to become features corresponding to object D. Therefore, when a user is interested in object A, object A in the first image needs to be identified to obtain relevant information about object A. Since some features of object A have been tampered with to become features corresponding to object D, directly identifying object A in the first image based on the tampered features of object A is equivalent to performing image recognition based on the features of object D. Consequently, the resulting recognition result will contain relevant information about object D, which may result in inaccurate image recognition results.
[0090] For example, while watching a video, a user becomes interested in a certain feature of a product from Brand A shown in the video and wants to identify products that have that feature. However, the trademark of the product in the video has been tampered with with that of Brand B, and Brand B's product does not have the feature shown in the video. If the user directly identifies the product image in the video, they will obtain information about Brand B's product. However, since Brand B's product does not have the feature shown in the video, the user will not be able to obtain information about products that have the feature shown in the video, resulting in the user not obtaining the identification result they were looking for.
[0091] Therefore, in order to make the image recognition results more accurate, it is necessary to determine and display the non-tampered areas in the first image, so that the user can distinguish between the tampered areas and the non-tampered areas in the first image, thereby avoiding the situation where the user obtains irrelevant information because he is not sure whether the image has been tampered with, thereby improving the accuracy of the recognition results.
[0092] Regarding the method of determining the non-tampered area that has not been tampered with in the first image, in a possible embodiment, a second image obtained by performing image enhancement on the first image can be obtained; the first image and the second image are input into a deep learning dual-stream detection model to obtain the output result of the deep learning dual-stream detection model, and the output result includes the confidence of the candidate tampered area in the first image; the candidate tampered area whose confidence is not less than a preset confidence threshold is determined as the tampered area in the first image.
[0093] Specifically, image enhancement methods may include image decoding, and / or image scaling, and / or spatial domain high-pass filtering, and / or frequency domain high-pass filtering. If the acquired first image is encoded and encapsulated image data, the first image needs to be decoded. Processing the first image through spatial domain high-pass filtering and / or frequency domain high-pass filtering can highlight the detailed features of the first image and enhance blurred edges in the first image, thereby achieving image enhancement.
[0094] Since some image features in the first image are relatively blurred, the accuracy of the image features extracted based on the first image is low. If only the first image is input into the deep learning dual-stream detection model to detect the tampered area, the accuracy of the image features extracted by the deep learning dual-stream detection model will be low, thereby making the accuracy of the output result of the deep learning dual-stream detection model low, and further making the accuracy of the non-tampered area in the determined first image low.
[0095] Since the second image is obtained by enhancing the first image, some of the image features of the second image are altered and different from those of the first image. Therefore, if only the second image is input into the deep learning dual-stream detection model to detect tampered areas, the image features extracted by the deep learning dual-stream detection model may be different from those of the first image, resulting in lower accuracy of the image features extracted by the deep learning dual-stream detection model, and thus lower accuracy of the output results of the deep learning dual-stream detection model, and further lower accuracy of the non-tampered areas in the first image.
[0096] Therefore, it is necessary to input the second image and the first image into the deep learning dual-stream detection model so that the image features extracted by the deep learning dual-stream detection model are obtained by integrating the image features of the first image and the image features of the second image, so that the accuracy of the image features extracted by the deep learning dual-stream detection model is higher, thereby making the accuracy of the output results of the deep learning dual-stream detection model higher, and further making the accuracy of the non-tampered area in the determined first image higher, thereby improving the accuracy of image recognition.
[0097] It is understandable that, for the first image, the tampered area and the non-tampered area are complementary. Once the tampered area in the first image is determined, the remaining areas in the first image other than the tampered area are the non-tampered areas in the first image that have not been tampered with. Therefore, determining the tampered area in the first image is equivalent to determining the non-tampered area in the first image that has not been tampered with.
[0098] If the confidence level of the candidate tampered region is not less than a preset confidence threshold, the candidate tampered region is determined as the tampered region in the first image, and the rest of the first image except the tampered region is the non-tampered region in the first image. The preset confidence threshold can be set based on past experience or actual needs, for example, the preset confidence threshold can be 0.5, 0.6, 0.65, etc.
[0099] This embodiment is selected because the deep learning dual-stream detection model (Faster-RCNN) has a faster detection speed and higher detection result accuracy compared to other RCNN detection models. Therefore, the deep learning dual-stream detection model can be used to improve the rate and accuracy of determining the tampered area in the first image. As described above, determining the tampered area in the first image is equivalent to determining the non-tampered area in the first image that has not been tampered with. Therefore, the deep learning dual-stream detection model can be used to improve the rate and accuracy of determining the non-tampered area in the first image, thereby improving the accuracy of image recognition.
[0100] In another possible embodiment, the non-tampered area in the first image may also be determined by using other detection models other than the deep learning dual-stream detection model.
[0101] In S302, when displaying the first image marked with the non-tampered area, it is necessary to make the non-tampered area and the tampered area in the displayed first image clearly distinguishable, so as to avoid the situation where the user obtains irrelevant information due to uncertainty about whether the image has been tampered.
[0102] In order to enable the user to more clearly distinguish between the non-tampered area and the tampered area in the first image, in a possible embodiment, the tampered area in the first image can be processed to obtain a third image; the image processing includes: framing the tampered area, mosaic processing, coloring processing, and binarization processing; and displaying the third image.
[0103] Specifically, based on the determination of the tampered area in the first image in S301, image processing can be performed on the tampered area, and the image obtained after processing can be used as a third image to display the third image. Performing image processing on the tampered area in the first image is equivalent to marking the tampered area in the first image, and the tampered area and the non-tampered area in the first image are complementary. Therefore, marking the tampered area in the first image is equivalent to marking the non-tampered area in the first image, that is, performing image processing on the tampered area in the first image is equivalent to marking the non-tampered area in the first image. Therefore, displaying the third image obtained by performing image processing on the tampered area in the first image is equivalent to displaying the first image marked with the non-tampered area.
[0104] The image processing used should be image processing that can significantly change the visual effect of the image. Therefore, the image processing may include: using a frame of any color, line or shape to select the tampered area in the first image, mosaic processing, coloring processing, binarization processing, etc. This application does not impose any restrictions on the specific method of image processing.
[0105] Exemplarily, the displayed first image marked with the non-tampered area can be as shown in Figure 4a. The dotted box in Figure 4a selects the tampered area in the first image, and the area not selected by the dotted box is the non-tampered area in the first image.
[0106] By selecting this embodiment, a third image obtained by performing image processing on the tampered area in the first image can be displayed, so that the first image marked with the non-tampered area can be displayed, so that the non-tampered area and the tampered area in the displayed first image can be clearly distinguished, thereby avoiding the situation where the user obtains irrelevant information because he is not sure whether the image has been tampered with, thereby improving the accuracy of the recognition result.
[0107] In S303, after displaying the first image with the non-tampered area marked, the user can perform a region selection operation on the non-tampered area. The region to be identified selected by the region selection operation is the region in the first image selected by the user where the object of interest is located. The region to be identified selected by the region selection operation can include only the non-tampered area, only the tampered area, or both the non-tampered area and the tampered area.
[0108] It is understandable that the image after the tampered area in the first image is tampered with may be a real image, such as an image obtained by cutting out other images. Therefore, the user may also be interested in the content in the tampered area, and the area to be identified may also include the tampered area.
[0109] Exemplarily, the first image for determining the area to be identified may be as shown in FIG4 b , where the area to be identified is framed by a solid line.
[0110] In S304, because the area to be recognized may contain multiple identifiable objects, and different objects also have multiple identifiable features, if recognition is performed directly on the area to be recognized, the recognition results obtained are relatively broad and may not include the information the user is interested in. For example, if the content in the area to be recognized that the user is interested in is a person wearing red clothes, if recognition is performed directly on the area to be recognized, the person will be recognized directly, resulting in a variety of recognition results such as a person sitting down and a person running, which may not include the person wearing red clothes that the user is interested in, resulting in low image recognition accuracy.
[0111] Based on this, the user can determine the target search attributes for the area to be identified in the preset search attributes according to the content of interest through S304-S306, thereby identifying the area to be identified according to the target search attributes and obtaining the identification result of interest to the user.
[0112] Preset search attributes can include a variety of commonly used identification object features. For example, the preset search attributes may include human features, such as a mannequin, hair, and walking posture; vehicle features, such as a license plate number and vehicle appearance; and object features, such as clothing style, teacup shape, and ornament shape. This application does not impose any restrictions on the specific settings of the preset search attributes.
[0113] When determining candidate search attributes, since there may be multiple identifiable objects in the area to be identified, if the categories of the multiple objects in the area to be identified are similar, the intersection of the preset search attributes possessed by each object in the area to be identified can be used as the candidate search attribute. For example, assuming that the area to be identified includes multiple different varieties of flowers, and the categories of the multiple objects in the area to be identified are all plants, the intersection of the preset search attributes possessed by each object, that is, the preset search attributes possessed by all flowers and plants, such as petal shape, leaf shape, etc., can be used as the candidate search attribute.
[0114] If the categories of multiple objects in the area to be identified are quite different, the union of the preset search attributes possessed by each object in the area to be identified can be used as an alternative search attribute. For example, if the area to be identified contains multiple objects such as people and vehicles, it is obvious that people and vehicles belong to different categories, and the categories to which people belong are quite different from the categories to which vehicles belong. There is almost no intersection between the preset search attributes possessed by people and the preset search attributes possessed by vehicles. Therefore, it is necessary to use the union of the preset search attributes possessed by each object, that is, the preset search attributes possessed by people and the preset search attributes possessed by vehicles as alternative search attributes, such as: human body model, hair, walking posture, license plate number, vehicle appearance, etc.
[0115] In S305, displaying an attribute selection interface including alternative search attributes may mean that the displayed attribute selection interface only includes alternative search attributes, or may mean that in the displayed attribute selection interface, the alternative search attributes are placed in front of other preset search attributes that are not alternative search attributes.
[0116] Regarding the display method of the attribute selection interface, in one possible implementation, after determining the area to be identified selected by the user's area selection operation, the alternative search attributes can be directly determined according to the method in S304, and the attribute selection interface containing the alternative search attributes can be displayed.
[0117] In another possible embodiment, it may also be in response to the interface display operation for the area to be identified, and the attribute selection interface containing the alternative retrieval attributes selected by the interface display operation may be displayed. Exemplarily, the display process of the attribute selection interface may be as shown in Figure 4c, where the retrieval type in Figure 4c is the preset retrieval attribute in this article, and the retrieval result is the recognition result in this article. As shown in Figure 4c, the attribute selection interface may be that after the area to be identified is determined, the user clicks on the retrieval type in the lower left corner of the first image, and the interface of the first image marked with the non-tampered area pops up in the form of a drop-down box, thereby displaying the attribute selection interface. The attribute selection interface may also be that after the area to be identified is determined, the user clicks on the retrieval type, and another interface pops up for the user to select the target retrieval attribute. In this example, the operation of the user clicking on the retrieval type can be regarded as an interface display operation for the area to be identified. This application does not impose any restrictions on the display form of the attribute selection interface.
[0118] As shown in Figure 4c, the alternative search attributes displayed in the attribute selection interface can include: human body model (i.e., human body comparison), walking posture, and clothing color. When the alternative search attribute is human body model, the human body model obtained by modeling the area to be identified is compared with the human body models stored in the database. The human body model in the database that is most similar to the human body model obtained based on the area to be identified is determined as the recognition result. This process is equivalent to a human body comparison. Therefore, the alternative search attribute of human body model corresponds to the human body comparison in Figure 4c.
[0119] In S306 , the target retrieval attribute may be regarded as a feature of an object in the area to be identified, and the number of the determined target retrieval attributes may be one or more.
[0120] According to the description in the aforementioned S304, the intersection or union of the preset retrieval attributes possessed by each object in the area to be identified can be used as an alternative retrieval attribute. If the union of the preset retrieval attributes possessed by each object in the area to be identified is used as an alternative retrieval attribute, the number of alternative retrieval attributes may be large, and the user cannot directly determine the alternative retrieval attributes possessed by the content he is interested in as the target retrieval attribute. Therefore, in a possible implementation, the user can determine the target retrieval attribute to be selected from the alternative retrieval attributes by text search in the attribute selection interface. In this example, the text search that the user can perform in the attribute selection interface can be used as the attribute selection operation of the user for the attribute selection interface, and the alternative retrieval attributes obtained by the user's search can be used as the alternative retrieval attributes selected by the attribute selection operation.
[0121] If the number of candidate search attributes is small, the user can directly identify the candidate search attributes possessed by the content they are interested in as the target search attribute. Therefore, in another possible implementation, the user can directly identify the target search attribute from the candidate search attributes. As shown in FIG4c , the user can select a target search attribute by checking a checkbox in front of the target search attribute displayed in the attribute selection interface as the user's attribute selection operation on the attribute selection interface. In response to the attribute selection operation, the candidate search attribute selected by the attribute selection operation can be identified as the target search attribute.
[0122] In S307, the attribute value can be regarded as the characteristic value of the target retrieval attribute, that is, the attribute value represents the characteristic value of the target retrieval attribute of an object in the area to be identified. The attribute value can be in various forms such as a characteristic vector and a numerical value.
[0123] For example, if the determined target retrieval attribute is vehicle appearance, it is necessary to determine the characteristic value of the vehicle appearance of the vehicle in the area to be identified; if the determined target retrieval attribute is animal appearance, it is necessary to determine the characteristic value of the animal appearance of the animal in the area to be identified.
[0124] In a possible embodiment, if the first image is a video frame in a video stream; if the target retrieval attribute is walking posture, the attribute value of the walking posture of the area to be identified is determined based on the first image and multiple video frames adjacent to the first image in the video stream.
[0125] In this embodiment, when the determined target retrieval attribute is a feature that cannot be reflected by a single image, the attribute value of the target retrieval attribute of the area to be identified can be determined through multiple video frames adjacent to the first image in the video stream to which the first image belongs and the first image.
[0126] In S308, the database is established based on the attribute values of the preset retrieval attributes of each object in the database, and the target retrieval attributes for the area to be identified are determined from the preset retrieval attributes. Therefore, when performing image recognition, the objects in the database can be screened by finding objects whose attribute values of the target retrieval attributes match the target attribute values.
[0127] In a possible embodiment, for each object in the database, the target attribute value and the attribute value of the target retrieval attribute of the object can be input into the deep learning network model to obtain an output result of whether the deep learning network model outputs a match, thereby determining the object in the database whose attribute value of the target retrieval attribute matches the target attribute value as the recognition result.
[0128] In another possible embodiment, the matching degree between the target attribute value and the attribute value of the target retrieval attribute of each object in the database can be determined as the matching degree of each object in the database; and the objects with a matching degree higher than a preset matching degree threshold can be determined as the recognition result.
[0129] The matching degree can be measured by similarity. Specifically, it can be measured by parameters negatively correlated with similarity, such as Euclidean distance and Mahalanobis distance, or by parameters positively correlated with similarity, such as cosine distance.
[0130] It can be understood that when determining an object with a matching degree higher than a preset matching degree threshold as a recognition result, it is not essentially a comparison of the specific numerical values of the matching degree, but a comparison of the matching degree between the image data in the area to be recognized and the objects in the database, so as to obtain an object with a higher matching degree between the image data in the area to be recognized as the recognition result.
[0131] If the matching degree is measured by the Euclidean distance which is negatively correlated with the similarity, the larger the Euclidean distance of the object, the smaller the similarity, and the lower the matching degree between the object and the image data in the area to be identified. At this time, it is necessary to determine the object whose Euclidean distance is less than the preset Euclidean distance threshold as the recognition result, and obtain the object with a higher matching degree with the image data in the area to be identified as the recognition result. Therefore, in this case, determining the object whose Euclidean distance is less than the preset Euclidean distance threshold as the recognition result is equivalent to determining the object whose matching degree is higher than the preset matching degree threshold as the recognition result.
[0132] The preset matching threshold can be set based on past experience or actual needs. If matching is measured by similarity, the preset matching threshold can be 0.5, 0.6, 0.65, etc. If matching is measured by Euclidean distance, the preset matching threshold can be 10, 15, 20, etc. As described above, in this case, objects with a Euclidean distance less than the preset matching threshold are determined as recognition results.
[0133] For example, if the target retrieval attribute is a human body model, the area to be identified can be input into the deep learning network to obtain a human body model A of the area to be identified output by the deep learning network; and the degree of matching between the human body model and each human body model in the database is determined.
[0134] In this example, human body model A is a specific human body model modeled based on the region to be identified. Human body model A is the attribute value of the target retrieval attribute of the region to be identified. Each human body model in the database is the attribute value of the target retrieval attribute of each object in the database.
[0135] Using a preset match threshold to filter objects is intended to achieve more accurate recognition results. However, the preset match threshold may be set too high or too low. If the preset match threshold is too high, it may result in an inability to determine the recognition result, which will lead to recognition failure and a lower recognition success rate. If the preset match threshold is too low, the recognition results may contain objects with low matches and inaccurate results, which may reduce the accuracy of the recognition results.
[0136] Based on this, in a possible embodiment, objects whose matching degrees are in the first preset number of places in the order of matching degrees from high to low may be determined as recognition results.
[0137] Specifically, instead of sorting, objects can be compared with each other's matching scores to determine the objects with matching scores that are in the first preset number of places in the matching score sorting from high to low as the recognition result. Alternatively, the matching scores can be sorted from high to low, and then the objects with matching scores that are in the first preset number of places in the sorting can be determined as the recognition result. Alternatively, the matching scores can be sorted from low to high, and then the objects with matching scores that are in the last preset number of places in the sorting can be determined as the recognition result. The preset number can be set according to user needs, for example, the preset number can be 5, 10, 15, etc.
[0138] By selecting this embodiment, a preset number of objects with higher matching degrees can be determined as recognition results based on the matching degrees of each object, so that the matching degrees of the objects included in the recognition results are higher than the matching degrees of other objects in the database, and the preset number can be set according to user needs, so that the number of objects included in the recognition results meets the user's needs and improves the accuracy of image recognition.
[0139] When there are multiple target retrieval attributes to be determined, the matching degree between the attribute value of the target retrieval attribute of the area to be identified and the attribute value of the target retrieval attribute of each object in the database can be determined for each target retrieval attribute, as the matching degree of each object in the database; for each target retrieval attribute, the ranking of the matching degree of each object in the matching degree ranking from high to low is determined; the ranking of each object under each target retrieval attribute is comprehensively considered to determine the comprehensive ranking of each object, and the objects in the first preset number of places are used as the recognition results.
[0140] For example, if the determined target retrieval attributes are target retrieval attribute 1 and target retrieval attribute 2, the matching degree between the attribute value of target retrieval attribute 1 of object 1 in the database and the attribute value of target retrieval attribute 1 of the area to be identified is ranked 12th, and the matching degree between the attribute value of target retrieval attribute 2 of object 1 in the database and the attribute value of target retrieval attribute 2 of the area to be identified is ranked 10th, then the comprehensive ranking of object 1 is (12+10)÷2=11. The matching degree between the attribute value of target retrieval attribute 1 of object 2 in the database and the attribute value of target retrieval attribute 1 of the area to be identified is ranked 3rd, and the matching degree between the attribute value of target retrieval attribute 2 of object 2 in the database and the attribute value of target retrieval attribute 2 of the area to be identified is ranked 7th, then the comprehensive ranking of object 1 is (3+7)÷2=5.
[0141] When calculating the comprehensive ranking of each object, you can take the average value as in the above example, or you can perform weighted summation according to the weight of each target retrieval attribute. This application does not impose any restrictions on the method of calculating the comprehensive ranking of each object.
[0142] For the display of recognition results, the thumbnail information of each recognition result can be displayed in sequence according to the order from high to low according to the matching degree between the attribute value of the target retrieval attribute of the recognition result and the target attribute value; in response to the viewing operation for the thumbnail information, the recognition result corresponding to the thumbnail information selected by the viewing operation is identified; and the detailed information of the recognition result is displayed.
[0143] In this example, the recognition result is displayed in the form of thumbnail information. The user can click on the displayed thumbnail information to view detailed information such as pictures, videos, or text corresponding to the recognition result. The operation of clicking on the displayed thumbnail information can be regarded as a viewing operation for the thumbnail information.
[0144] Exemplarily, the display of the recognition results can be shown as the search results on the right side of Figure 4c, where NO1, NO2, NO3, NO4, and NO5 refer to the order in which the objects in the recognition results are arranged from high to low according to the matching degree, that is, the matching degree of the object displayed at NO1 is the highest in the recognition results.
[0145] After determining the non-tampered area in the first image that has not been tampered with in the aforementioned step S301, the first image with the non-tampered area marked and the first image without the non-tampered area marked can be spliced and displayed for user viewing. Therefore, in a possible embodiment, if the first image is a single image, two spliced first images can be displayed, wherein one first image is marked with the non-tampered area and the other first image is not marked with the non-tampered area.
[0146] If the first image is the video frames in a video clip, two video playback windows spliced together can be displayed. The two video playback windows play the video clip synchronously, and the video frames played by one of the video playback windows are marked with non-tampered areas, while the video frames played by the other video playback window are not marked with non-tampered areas.
[0147] It can be understood that the aforementioned S301-S302 can be regarded as retrieving the non-tampered area in the first image, and S303-S308 can be regarded as retrieving the object in the area to be identified. Therefore, the process of retrieving the area to be identified can be regarded as the process of the second retrieval of the first image. For the convenience of description below, the process of retrieving the area to be identified will be referred to as the secondary retrieval.
[0148] FIG5 is another flow chart of an image recognition method provided by an embodiment of the present application. The method shown in FIG5 can be used to determine the tampered area in the first image. As shown in FIG5, the method includes:
[0149] S501: Input a first image.
[0150] S502: Perform image preprocessing on the first image.
[0151] S503 , obtaining an original image (ie, a first image) through image preprocessing and a noise image (ie, the aforementioned second image) obtained through preprocessing of spatial domain high-pass filtering and frequency domain high-pass filtering.
[0152] S504: Input the first image and the second image into a deep learning dual-stream detection model (Faster-RCNN).
[0153] S505, the deep learning dual-stream detection model outputs the tampered image area (i.e., the aforementioned candidate tampered area) and the confidence of the candidate tampered area.
[0154] S506, determine whether the confidence of the candidate tampering area is greater than the preset confidence threshold through the alarm unit. If not, that is, the confidence of the candidate tampering area is not greater than the preset confidence threshold, execute S507; if yes, that is, the confidence of the candidate tampering area is greater than the preset confidence threshold, execute S508.
[0155] S507, the process ends.
[0156] S508 , converting the normalized coordinates of the candidate tampered region output by the model into the actual position of the candidate tampered region in the first image, thereby achieving coordinate conversion.
[0157] S509, drawing the tampered area according to the actual position of the candidate tampered area in the first image, that is, performing image processing on the tampered area in the first image in the aforementioned S302, realizes the labeling of the tampered area in the first image, which is equivalent to realizing the labeling of the non-tampered area in the first image.
[0158] S510: Splice and display the first image marked with the tampered area and the original first image.
[0159] For example, as shown in Figure 6, the left side of Figure 6 shows the original first image, and the right side shows the first image with the tampered area marked. The dotted box selects the tampered area in the first image. Therefore, the spliced image shown in Figure 6 can be regarded as the original image and the tampered display image.
[0160] S511, determining whether the input first image is a video formed by multiple related first images arranged in time sequence, if not, executing step S512; if yes, executing S513 and S514.
[0161] S512, the process ends.
[0162] S513, synchronize the timestamp to control the playback display.
[0163] S514, the process ends.
[0164] FIG7 is another flow chart of the image recognition method provided by an embodiment of the present application. As shown in FIG7 , based on the embodiment of FIG5 , after the tampered region is drawn according to the actual position of the candidate tampered region in the first image with a confidence greater than the confidence threshold, and the first image with the tampered region drawn is displayed in the secondary search interface, that is, after the non-tampered region in the first image is marked and the first image with the non-tampered region marked is displayed in the secondary search interface, step S701 is executed. The secondary search interface is a user interface for implementing the secondary search (i.e., the aforementioned S303-S308).
[0165] S701, secondary search configuration, the user selects the search area (ie the aforementioned area to be identified) and the search type (ie the aforementioned target search attribute).
[0166] The user can enter an area selection command in the secondary search interface to determine the area to be identified indicated by the area selection command. After determining the area to be identified, the user can enter an attribute selection command in the displayed attribute selection interface to determine the target search attribute indicated by the attribute selection command.
[0167] S702: Identify or model the area to be identified selected by the user through a deep learning structured / model processing module to obtain structured data or model data.
[0168] S703, searching through the data retrieval model, that is, searching the database through the structured data or model data to obtain the search results. The database may store structured data or models of various objects.
[0169] S704 , displaying the most similar top display according to the search results, that is, displaying the search results (ie, the aforementioned recognition results) that are most similar to the structured data or model data obtained by recognizing or modeling the area to be recognized.
[0170] S705: The user clicks on the entry information, and the detailed information of the recognition result is displayed through the detailed information display module.
[0171] When displaying the recognition results, brief information may be displayed. The user may click on the item information to view detailed information of the recognition results, and the detailed information may be displayed through the detailed information display module.
[0172] S706, the process ends.
[0173] The above steps are steps S303 to S308 performed on the basis of the embodiment of FIG. 5 .
[0174] For example, if a user needs to identify a person in a first image, but the area above the torso of the person has been tampered with, the specific implementation steps of the image recognition method provided by this application are as follows: input a first image; perform spatial domain high-pass filtering and frequency domain high-pass filtering on the first image to obtain a second image. Input the first image and the second image into a deep learning dual-stream detection model (Faster-RCNN), and obtain the area above the torso output by the deep learning dual-stream detection model and the confidence of the area above the torso; the alarm unit determines that the confidence of the area above the torso is greater than a preset confidence threshold, that is, the area above the torso of the person in the first image has been tampered with, then the normalized coordinates of the area above the torso output by the model are converted to the actual position of the area above the torso in the first image, and according to the actual position of the area above the torso in the first image, the area above the torso is framed out in the first image using a dotted box. The area above the torso is the area shown in the dotted box in Figure 4a.
[0175] The user selects the torso of a person in the first image as the area to be identified and selects a human body model as the target retrieval attribute. The area to be identified is input into the deep learning model processing module for modeling, resulting in a human body model of the area to be identified. The human body model of the area to be identified is then compared with human body model objects stored in a database. The human body model object in the database that is most similar to the human body model in the area to be identified is determined as the recognition result, and the recognition result is displayed. The displayed information may only contain rough information about the recognition result; the user can view detailed information about the recognition result by clicking on the entry information.
[0176] The image recognition method provided in this application can be applied to electronic devices such as DVR (Digital Video Recorder), NVR (Network Video Recorder) and central storage devices.
[0177] Corresponding to the aforementioned image recognition method, an embodiment of the present application further provides an image recognition system, the system comprising:
[0178] The business processing unit is configured to display the first image marked with the tampered area.
[0179] In a possible embodiment, the business processing unit is further used to respond to a playback instruction input for the display interface, obtain the image included in the multimedia data indicated by the playback instruction as the first image, and obtain the tampering area obtained by pre-tampering detection of the first image; and display the first image marked with the tampering area in the multimedia data playback window of the display interface.
[0180] In a possible embodiment, the display interface further includes: a multimedia data list, the multimedia data list including identifiers corresponding to a plurality of multimedia data;
[0181] The service processing unit is further configured to set a tampering tag for an identifier corresponding to the multimedia data in the multimedia data list if the multimedia data includes a tampered image;
[0182] The service processing unit is further configured to, in response to a tampering detection result that a first image displayed in the multimedia data playback window is tampered with, display a tampering annotation display control in the display interface; in response to a tampering annotation hiding operation on the tampering annotation display control, hide the annotation of the tampered area in the first image; and display the first image marked with the tampered area in the multimedia data playback window of the display interface, including: in response to the tampering annotation display operation on the tampering annotation display control, display the annotation of the tampered area in the first image;
[0183] The first images are each video frame in the video clip;
[0184] The service processing unit is further configured to perform tampering detection on each video frame in the order of each video frame in the video clip, determine the video frame whose tampering detection result is tampered among the video frames as the tampered video frame; determine the time zone corresponding to the tampered video frame in the time progress bar of the video clip, and set a tampering mark in the time zone;
[0185] The business processing unit is also used to perform tamper detection on each video frame in the order of each video frame in the video clip, and display the tamper detection progress of the video clip in the display interface; the tamper detection progress is the ratio of the number of video frames in the video clip that have undergone tamper detection to the number of all video frames included in the video clip.
[0186] In a possible embodiment, the system further includes:
[0187] A deep learning image authenticity processing unit is used to determine a non-tampered area in the first image that has not been tampered with.
[0188] Specifically, the method in which the deep learning image authenticity processing unit determines the non-tampered area in the first image that has not been tampered with is the same as the method in the aforementioned S301. Please refer to the relevant instructions of the aforementioned S301 and will not be repeated here.
[0189] The business processing unit is also used to display a first image marked with a non-tampered area; in response to an area selection operation on the first image, determine the area to be identified selected by the area selection operation in the first image; determine the alternative retrieval attributes possessed by the objects existing in the area to be identified in the preset retrieval attributes; display an attribute selection interface containing the alternative retrieval attributes; in response to an attribute selection operation on the attribute selection interface, identify the alternative retrieval attributes selected by the attribute selection operation as the target retrieval attribute; obtain the attribute value of the target retrieval attribute of the area to be identified as the target attribute value; in a database that pre-stores the attribute values of each object and the preset retrieval attributes of each object, search for objects whose attribute values of the target retrieval attributes match the target attribute values as the identification results.
[0190] Specifically, the business processing unit displays a first image marked with a non-tampered area; in response to an area selection instruction input for the non-tampered area, determines the area to be identified indicated by the area selection instruction in the first image; identifies the object in the area to be identified, and obtains the identification result in the same manner as in the aforementioned S302-S308. Please refer to the relevant instructions of the aforementioned S302-S308, which will not be repeated here.
[0191] By selecting this embodiment, the non-tampered area in the first image that has not been tampered with can be determined through the deep learning image authenticity processing unit, and the first image marked with the non-tampered area can be displayed through the business processing unit. In response to the area selection operation on the first image, the area to be identified selected by the area selection operation is determined in the first image; the alternative retrieval attributes possessed by the objects existing in the area to be identified are determined in the preset retrieval attributes; an attribute selection interface containing the alternative retrieval attributes is displayed; in response to the attribute selection operation on the attribute selection interface, the alternative retrieval attributes selected by the attribute selection operation are identified as the target retrieval attributes, and the attribute value of the target retrieval attribute of the area to be identified is obtained as the target attribute value. In a database that pre-stores the attribute values of each object and the preset retrieval attributes of each object, objects whose attribute values of the target retrieval attributes match the target attribute values are searched for as recognition results. When determining the target retrieval attribute, the alternative retrieval attributes possessed by the objects present in the area to be identified can be determined from the preset retrieval attributes. Since the alternative retrieval attributes can be regarded as preset retrieval attributes that the user may select, the alternative retrieval attributes can be used as preset retrieval attributes recommended to the user for selection, thereby helping the user to quickly determine the target retrieval attributes possessed by the content that the user is interested in from multiple preset retrieval attributes. Since the database is established based on the attribute values of the preset retrieval attributes of each object, and the target retrieval attributes for the area to be identified are determined from the preset retrieval attributes, the target attribute value is the attribute value of the target retrieval attributes of the area to be identified. Therefore, by matching the attribute values of the target retrieval attributes of each object in the database with the target attribute values, each object in the database can be screened, and the objects whose attribute values of the target retrieval attributes in the database match the target attribute values are used as recognition results, thereby more accurately obtaining relevant information about the content that the user is interested in. Therefore, the embodiments of the present application can determine the characteristics of the content that the user is interested in by determining the target retrieval attributes, and by matching the attribute values of the target retrieval attributes of each object in the database with the target attribute values, more accurately determine relevant information about the content that the user is interested in, thereby improving the accuracy of image recognition.
[0192] In a possible embodiment, the system further includes: a data receiving unit, configured to receive the first image;
[0193] an image processing unit, configured to perform image enhancement on the first image to obtain a second image;
[0194] a deep learning image authenticity processing unit, specifically configured to obtain a second image; input the first image and the second image into a deep learning dual-stream detection model, obtain an output result of the deep learning dual-stream detection model, the output result including a confidence score of a candidate tampered region in the first image, and send the output result to an alarm unit;
[0195] an alarm unit, configured to, in response to the output result, determine a candidate tampered region in the output result whose confidence level is not less than a preset confidence threshold as a tampered region in the first image, and issue an alarm for the tampered region;
[0196] A storage management unit, used for storing data;
[0197] The configuration management unit is used to configure the image processing unit, the deep learning image authenticity processing unit, the alarm unit, the business processing unit and the storage management unit.
[0198] In a possible embodiment, the business processing unit is further configured to display the first image marked with a non-tampered area in response to an identification operation on the alarm unit.
[0199] For example, a schematic diagram of the structure of the image recognition system can be shown in Figure 8. After receiving the first image, the data receiving unit 801 inputs the data into the image processing unit (i.e., the picture / video processing unit 802), that is, the first image is input into the image processing unit 802, and the image processing unit 802 outputs a second image after pre-processing the first image. The image processing unit 802 can send the obtained second image to the storage management unit 806 for storage, or obtain the required image data from the storage management unit 806. The first image and the second image are input into the deep learning image authenticity processing unit 803, and the candidate tampering region in the first image and the authenticity confidence of the candidate tampering region output by the deep learning image authenticity processing unit 803 are obtained. The alarm unit 804 determines whether the candidate tampering region is the tampering region in the first image based on the authenticity confidence of the candidate tampering region, and issues an alarm for the tampering region to obtain an alarm result, thereby realizing the detection of the tampering region of the first image, that is, realizing the first recognition of the first image.
[0200] Based on the alarm result, the business processing unit 805 can display the first image marked with the non-tampered area, and use a pop-up window or voice broadcast to implement an alarm for the tampered area in the first image. When receiving the recognition instruction input by the user, that is, the user wants to perform a second recognition on the first image, it executes the aforementioned S302-S308 to perform image recognition on the first image. The business processing unit is the business processing 805 shown in Figure 8, and is hereinafter referred to as the business processing unit 805.
[0201] In addition, the business processing unit 805 can send the retrieval result, that is, the image recognition result, to the storage management unit 806, so that the storage management unit 806 saves the image recognition result. The alarm result obtained by the alarm unit 804 can also be sent to the storage management unit 806, so that the storage management unit 806 saves the alarm result.
[0202] Configuration management unit 807 is used to configure image processing unit 802, deep learning image authenticity processing unit 803, alarm unit 804, business processing unit 805, and storage management unit 806. For example, deep learning image authenticity processing unit 803 can be configured to output the authenticity confidence level of the candidate tampered area and determine whether to issue an alarm display based on the confidence level; alarm unit 804 can be configured to determine whether to store the alarm; and the area to be identified and the target retrieval attributes when business processing unit 805 executes steps S302-S304 can be configured. Configuration management unit 807 is the configuration management unit 807 shown in Figure 8.
[0203] FIG9 is another flow chart of the image recognition method provided by the present application. As shown in FIG9 , the configuration management unit configures the data receiving unit and determines different pre-processing methods for the first image received by the data receiving unit. The video processing unit (i.e., the aforementioned image processing unit) determines whether the received first image is an analog video. If the first image is an analog video, the first image is captured. If the first image is a digital video, the first image is decoded. If the first image is not a video, the image data of the first image is obtained. The second image obtained after processing by the image processing unit is sent to the storage management unit for storage, and the first image and the second image are input to the deep learning authenticity processing unit (i.e., the aforementioned deep learning image authenticity processing unit) to detect the tampered area in the first image, and obtain the candidate tampered area in the first image output by the deep learning authenticity processing unit and the confidence of the candidate tampered area. The alarm unit determines whether the candidate tampered area is the tampered area in the first image based on the authenticity confidence of the candidate tampered area, obtains an alarm result, and sends the alarm result to the storage management unit for storage. The business processing unit (i.e., business processing) can display the alarm result according to the alarm result. In addition, the business processing unit can display the first image marked with the non-tampered area based on the alarm result, and perform a secondary identification of the non-tampered area based on the object of the database in the storage management unit to obtain the identification result, that is, realize the secondary retrieval of the image, and display the identification result.
[0204] Corresponding to the aforementioned image recognition method, an embodiment of the present application further provides an image recognition device, as shown in FIG10 , comprising:
[0205] The tampered area display module 1001 is used to display the first image marked with the tampered area.
[0206] In a possible embodiment, displaying a first image marked with a tampered area includes: in response to a play instruction input on a display interface, obtaining an image included in multimedia data indicated by the play instruction as the first image, and obtaining the tampered area obtained by pre-tampering detection on the first image;
[0207] The first image marked with the tampered area is displayed in the multimedia data playback window of the display interface.
[0208] In a possible embodiment, the display interface further includes: a multimedia data list, wherein the multimedia data list includes identifiers corresponding to a plurality of multimedia data; and the device further includes:
[0209] The tampering label setting module is used to set a tampering label for the identifier corresponding to the multimedia data in the multimedia data list if the multimedia data includes a tampered image.
[0210] In a possible embodiment, the first image is each video frame in a video clip, and the apparatus further includes:
[0211] a tampered video frame determination module, configured to perform tampering detection on each video frame in the order of each video frame in the video clip, and determine the video frame whose tampering detection result is tampered among the video frames as the tampered video frame;
[0212] The tampering mark setting module is used to determine the time area corresponding to the tampered video frame in the time progress bar of the video segment and set the tampering mark in the time area.
[0213] In a possible embodiment, the device further includes:
[0214] a tampering annotation display control display module, configured to display a tampering annotation display control in a display interface in response to a tampering detection result indicating that the first image displayed in the multimedia data playback window has been tampered with;
[0215] A tampering annotation hiding module is used to hide the annotation of the tampered area in the first image in response to a tampering annotation hiding operation on a tampering annotation display control; or to display the first image marked with the tampering area in the multimedia data playback window of the display interface, including: displaying the annotation of the tampered area in the first image in response to a tampering annotation display operation on a tampering annotation display control.
[0216] In a possible embodiment, the first image is each video frame in a video clip, and the apparatus further includes:
[0217] The tampering detection progress display module is used to perform tampering detection on each video frame in the order of each video frame in the video clip, and to display the tampering detection progress of the video clip in the display interface; the tampering detection progress is the ratio of the number of video frames in the video clip that have been tampered with to the number of all video frames included in the video clip.
[0218] In a possible embodiment, the device further includes:
[0219] a non-tampered area determination module, configured to determine a non-tampered area that has not been tampered with in the first image;
[0220] A first image display module, configured to display the first image marked with a non-tampered area;
[0221] a module for determining an area to be identified, configured to determine, in response to an area selection operation on the first image, an area to be identified selected by the area selection operation in the first image;
[0222] The candidate retrieval attribute determination module is used to determine the candidate retrieval attributes possessed by the objects existing in the area to be identified from the preset retrieval attributes;
[0223] The attribute selection interface display module is used to display the attribute selection interface containing the alternative search attributes;
[0224] a target retrieval attribute selection module, configured to respond to an attribute selection operation on the attribute selection interface and identify a candidate retrieval attribute selected by the attribute selection operation as a target retrieval attribute;
[0225] A target attribute value acquisition module is used to acquire the attribute value of the target retrieval attribute of the area to be identified as the target attribute value;
[0226] The recognition result search module is used to search for objects whose attribute values of target search attributes match the target attribute values in a database pre-stored with various objects and attribute values of preset search attributes of various objects, as recognition results.
[0227] In a possible embodiment, the first image is a video frame in a video stream, and the target retrieval attribute is a walking posture;
[0228] The target attribute value acquisition module acquires the attribute value of the target retrieval attribute of the area to be identified as the target attribute value, including:
[0229] According to the first image and a plurality of video frames adjacent to the first image in the video stream, an attribute value of the walking posture of the area to be identified is determined as a target attribute value.
[0230] In a possible embodiment, the device further includes:
[0231] The abbreviated information display module is used to display the abbreviated information of each recognition result in order of the matching degree between the attribute value of the target retrieval attribute of the recognition result and the target attribute value from high to low;
[0232] a thumbnail information recognition module, configured to, in response to a viewing operation on the thumbnail information, identify a recognition result corresponding to the thumbnail information selected by the viewing operation;
[0233] The detailed information display module is used to display the detailed information of the recognition results.
[0234] In a possible embodiment, the non-tampering area determining module determines the non-tampering area that has not been tampered with in the first image, including:
[0235] Acquire a second image obtained by performing image enhancement on the first image;
[0236] Inputting the first image and the second image into a deep learning two-stream detection model to obtain an output result of the deep learning two-stream detection model, the output result including a confidence score of a candidate tampering region in the first image;
[0237] The candidate tampered region whose confidence level is not less than a preset confidence threshold is determined as the tampered region in the first image.
[0238] In a possible embodiment, the first image is a single image, and the device further includes:
[0239] A first splicing display module is used to display two first images spliced together, wherein one first image is marked with a non-tampered area, and the other first image is not marked with a non-tampered area;
[0240] or,
[0241] The first image is each video frame in the video clip, and the apparatus further includes:
[0242] The second splicing display module is used to display two video playback windows spliced together. The two video playback windows play video clips synchronously, and the video frame played by one of the video playback windows is marked with a non-tampered area, while the video frame played by the other video playback window is not marked with a non-tampered area.
[0243] In a possible embodiment, the first image display module displays the first image marked with the non-tampered area, including:
[0244] Performing image processing on the tampered area in the first image to obtain a third image; wherein the image processing includes: selecting the tampered area, mosaic processing, coloring processing, and binarization processing;
[0245] Display the third image.
[0246] The present application also provides an electronic device, as shown in FIG11 , including:
[0247] Memory 1101, used for storing computer programs;
[0248] The processor 1102 is configured to implement the following steps when executing the program stored in the memory 1101:
[0249] The first image is displayed with the tampered area marked.
[0250] Furthermore, the electronic device may further include a communication bus and / or a communication interface, and the processor 1102 , the communication interface, and the memory 1101 communicate with each other via the communication bus.
[0251] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0252] The communication interface is used for communication between the above electronic device and other devices.
[0253] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0254] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0255] In another embodiment provided in the present application, a computer-readable storage medium is further provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of any of the above-mentioned image recognition methods are implemented.
[0256] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any one of the image recognition methods in the above embodiments.
[0257] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a solid-state drive (SSD).
[0258] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0259] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system, device, electronic device, computer-readable storage medium, and computer program product embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For related portions, reference can be made to the descriptions of the method embodiments.
[0260] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An image display method, characterized in that: The method comprises: The first image is shown with the tampered areas annotated.
2. The method according to claim 1, characterized in that The displaying of the first image with the tampered area marked includes: In response to a play instruction input to a display interface, an image included in the multimedia data indicated by the play instruction is obtained as a first image, and a tampered area obtained by pre-tampering detection of the first image is obtained; The first image marked with the tampered area is displayed in the multimedia data playback window of the display interface.
3. The method according to claim 2, characterized in that The display interface further includes: a multimedia data list, wherein the multimedia data list includes identifiers corresponding to a plurality of multimedia data; the method further includes: If the multimedia data includes a tampered image, a tampered tag is set for the identifier corresponding to the multimedia data in the multimedia data list.
4. The method according to claim 2, characterized in that: The first images are video frames in a video clip, and the method further includes: According to the order of the video frames in the video clip, each video frame is subjected to tampering detection, and a tampered video frame is determined as a tampered video frame as a tampered video frame; The time zone corresponding to the tampered video frame in the time progress bar of the video segment is determined, and a tampering mark is set in the time zone.
5. The method according to claim 2, characterized in that: The method further comprises: In response to a tampering detection result that the first image displayed in the multimedia data playback window is tampered with, displaying a tampering mark display control in the display interface; In response to a tampering annotation hiding operation on the tampering annotation display control, hiding the annotation of the tampered area in the first image; or, displaying the first image marked with the tampered area in the multimedia data playback window of the display interface, including: in response to a tampering annotation display operation on the tampering annotation display control, displaying the annotation of the tampered area in the first image.
6. The method according to claim 2, characterized in that The first images are video frames in a video clip, and the method further includes: According to the order of each video frame in the video clip, each video frame is checked for tampering, and the progress of tampering detection on the video clip is displayed in the display interface; the tampering detection progress is the ratio of the number of video frames in the video clip that have been checked for tampering to the number of all video frames included in the video clip.
7. The method according to claim 1, characterized in that The method further comprises: determining a non-tampered region in the first image that has not been tampered with; displaying the first image with the non-tampered area marked; In response to a region selection operation on the first image, determining, in the first image, a region to be identified selected by the region selection operation; Determining, from the preset search attributes, candidate search attributes possessed by the object existing in the to-be-identified area; Displaying an attribute selection interface including the candidate search attributes; In response to an attribute selection operation on the attribute selection interface, identifying a candidate search attribute selected by the attribute selection operation as a target search attribute; Acquire the attribute value of the target retrieval attribute of the to-be-identified area as the target attribute value; In a database pre-stored with each object and the attribute value of the preset search attribute of each object, an object whose attribute value of the target search attribute matches the target attribute value is searched as a recognition result.
8. The method according to claim 7, characterized in that The first image is a video frame in a video stream, and the target retrieval attribute is a walking posture; The step of obtaining the attribute value of the target retrieval attribute of the area to be identified as the target attribute value includes: According to the first image and a plurality of video frames adjacent to the first image in the video stream, an attribute value of the walking posture of the area to be identified is determined as a target attribute value.
9. The method according to claim 7, characterized in that: The method further comprises: Displaying the abbreviated information of each recognition result in sequence according to the order of the matching degree between the attribute value of the target search attribute of the recognition result and the target attribute value from high to low; In response to a viewing operation on the thumbnail information, identifying a recognition result corresponding to the thumbnail information selected by the viewing operation; Displays detailed information of the recognition result.
10. The method according to claim 7, characterized in that Determining the non-tampered area in the first image that has not been tampered with includes: Acquire a second image obtained by performing image enhancement on the first image; Inputting the first image and the second image into a deep learning two-stream detection model to obtain an output result of the deep learning two-stream detection model, wherein the output result includes a confidence score of a candidate tampered region in the first image; The candidate tampered region whose confidence is not less than a preset confidence threshold is determined as the tampered region in the first image.
11. The method according to claim 7, characterized in that The first image is a single image, and the method further includes: Displaying two first images spliced together, wherein one first image is marked with the non-tampered area, and the other first image is not marked with the non-tampered area; or, The first images are video frames in a video clip, and the method further includes: Two video playback windows spliced together are displayed, the two video playback windows synchronously play the video clip, and the video frame played by one of the video playback windows is marked with the non-tampered area, while the video frame played by the other video playback window is not marked with the non-tampered area.
12. The method according to claim 7, characterized in that The displaying of the first image with the non-tampering area marked thereon includes: Performing image processing on the tampered area in the first image to obtain a third image; wherein the image processing includes: selecting the tampered area, mosaic processing, coloring processing, and binarization processing; The third image is displayed.
13. An image recognition system, characterized in that: The system comprises: The business processing unit is used to display the first image marked with the tampered area.
14. The system according to claim 13, characterized in that The service processing unit is further configured to, in response to a play instruction inputted to the display interface, obtain an image included in the multimedia data indicated by the play instruction as a first image, and obtain a tampered area obtained by pre-tampering detection of the first image; Displaying the first image with the tampered area marked in the multimedia data playback window of the display interface; The display interface further includes: a multimedia data list, wherein the multimedia data list includes identifiers corresponding to a plurality of multimedia data; The service processing unit is further configured to set a tampering tag for an identifier corresponding to the multimedia data in the multimedia data list if the multimedia data includes a tampered image; The business processing unit is further configured to, in response to a tampering detection result that the first image displayed in the multimedia data playback window is tampered with, display a tampering mark display control in the display interface; in response to a tampering mark hiding operation on the tampering mark display control, hide the mark of the tampered area in the first image; the displaying of the first image marked with the tampered area in the multimedia data playback window of the display interface includes: in response to the tampering mark display operation on the tampering mark display control, display the mark of the tampered area in the first image; The first images are video frames in the video clip; The business processing unit is further used to perform tampering detection on each video frame in the order of each video frame in the video clip, determine the video frame whose tampering detection result is tampered in each video frame as the tampered video frame; determine the time area corresponding to the tampered video frame in the time progress bar of the video clip, and set a tampering mark in the time area; The business processing unit is also used to perform tampering detection on each video frame in the video clip according to the order of each video frame in the video clip, and display the tampering detection progress of the video clip in the display interface; the tampering detection progress is the ratio of the number of video frames in the video clip that have undergone tampering detection to the number of all video frames included in the video clip.
15. The system according to claim 13, characterized in that The system also includes: a deep learning image authenticity processing unit; The deep learning image authenticity processing unit is used to determine a non-tampered area that has not been tampered with in the first image; The business processing unit is also used to display the first image marked with the non-tampered area; in response to the area selection operation on the first image, determine the area to be identified selected by the area selection operation in the first image; determine the alternative retrieval attributes possessed by the objects existing in the area to be identified in the preset retrieval attributes; display an attribute selection interface containing the alternative retrieval attributes; in response to the attribute selection operation on the attribute selection interface, identify the alternative retrieval attributes selected by the attribute selection operation as the target retrieval attribute; obtain the attribute value of the target retrieval attribute of the area to be identified as the target attribute value; in a database that pre-stores the attribute values of each object and the preset retrieval attributes of each object, search for objects whose attribute values of the target retrieval attributes match the target attribute values as the identification result.
16. The system according to claim 15, characterized in that The system further comprises: A data receiving unit, configured to receive the first image; An image processing unit, configured to perform image enhancement on the first image to obtain a second image; The deep learning image authenticity processing unit is specifically used to obtain the second image; input the first image and the second image into the deep learning dual-stream detection model, obtain the output result of the deep learning dual-stream detection model, the output result includes the confidence of the candidate tampering area in the first image, and send the output result to the alarm unit; The alarm unit is configured to, in response to the output result, determine the candidate tampered region in the output result whose confidence is not less than a preset confidence threshold as the tampered region in the first image, and issue an alarm for the tampered region; A storage management unit, used for storing data; A configuration management unit is used to configure the image processing unit, the deep learning image authenticity processing unit, the alarm unit, the business processing unit and the storage management unit.
17. The system according to claim 15, characterized in that The business processing unit is further configured to display the first image marked with the non-tampered area in response to an identification operation on the alarm unit.
18. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-12 when executing a program stored in a memory.
Citation Information
Patent Citations
Image tampering recognition method and device, server and storage medium
CN111415336A
Tamper identification method and device, computer equipment and storage medium
CN117115823A
Image recognition method and system and electronic equipment
CN117407562A
Method, device, and computer program for providing image search information
US20210326375A1