Target detection method and video processing equipment

By accepting text, channel and time period as search conditions and combining it with the ITR model for cross-modal retrieval, the problem of low efficiency of target detection in recorded videos is solved, and fast and accurate target detection is achieved.

CN120670622APending Publication Date: 2025-09-19HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510765856.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When performing target detection in a video, existing technologies require viewing the video frame by frame to obtain an image of the target object, resulting in low detection efficiency.

Method used

This paper provides a target detection method that accepts text, channel, and time period as search conditions, combines the ITR model for cross-modal retrieval, quickly finds images matching target features in recorded videos, and displays the detected target information.

Benefits of technology

The efficiency of target detection in recorded videos has been improved. Users can accurately search for targets without having to review the recorded videos frame by frame, which improves detection speed and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670622A_ABST
    Figure CN120670622A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a target detection method and video processing equipment, and relates to the technical field of image processing. The method comprises the steps of receiving input of a search condition, and displaying a first image of at least one first target hit by the search condition; receiving and responding to a search instruction used for indicating to search a first target in a certain first image, and respectively displaying target information of each second target detected in the first video; wherein the first image is from the first video, the first video is from the video of the first channel in the first time period, the first target is a target matched with the first text, and the second target is a target matched with the first target in the first image indicated by the image search instruction. By adopting the embodiment of the invention, the efficiency of target detection in the recorded video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a target detection method and video processing equipment. Background Art

[0002] With the increasing demand for security, target detection in video footage based on the target object's image has become a core technical approach in the public safety field. Related technologies typically begin by acquiring an image of the target object. Then, using a target detection algorithm, they extract the target object's appearance features, such as facial features, from the image. These features are then used as the target features. Candidate targets are then detected frame by frame in the video, generating candidate frames. The similarity between the appearance features of the candidate targets in the candidate frames and the target features is then calculated. Candidate targets that meet the similarity criteria are considered the target objects. The target object's location can then be determined based on the correspondence between the video frame containing the target object and the camera.

[0003] However, in the actual target detection process, it is difficult to obtain the image of the target object. For example, only the witness knows the facial features and clothing features of the thief. The witness needs to cooperate in viewing the video footage of the relevant time period and relevant location frame by frame to obtain the target object. The witness may even need to view the video footage of other irrelevant locations frame by frame, resulting in low efficiency of target detection based on video footage. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a target detection method and video processing device to improve the efficiency of target detection in recorded video. The specific technical solution is as follows:

[0005] In a first aspect of an embodiment of the present application, a target detection method is provided, the method comprising:

[0006] Receive input of search conditions and display a first image of at least one first target matched by the search conditions; wherein the search conditions include a first channel, a first time period, and a first text; the first image is from a first video; the first video is from a video of the first channel within the first time period; the first target is a target that matches the first text; and the first text includes: a subtext for indicating a target category and a subtext for indicating features possessed by the target;

[0007] An image search instruction is received, where the image search instruction is used to instruct a search for a first target in a first image, and, in response to the image search instruction, target information of each second target detected in the first video is displayed respectively, wherein the target information includes: a timestamp of a video frame indicating that the second target is detected and information of a channel from which the video frame originates, the second target being a target that matches the first target in the first image indicated by the image search instruction.

[0008] A second aspect of the embodiments of the present application provides a target detection method, the method comprising:

[0009] Receive input of search conditions and display a first image of at least one first target hit by the search conditions; wherein the search conditions include a first channel, a first time period and a first text, the first image comes from a first video, the first video comes from a video of the first channel within the first time period, the first target is a target that matches the first text, and the first text includes: a subtext for representing a target category and a subtext for representing characteristics possessed by the target.

[0010] According to a third aspect of the embodiments of the present application, a video processing device is provided, which is used to implement the target detection method as described in the first aspect.

[0011] Beneficial effects of the embodiments of the present application:

[0012] An embodiment of the present application provides a target detection method and video processing device, which combines text, channels, and time periods as search conditions. In this way, when searching for a target, the user only needs to enter text that matches the characteristics of the target to be found and the channel and time period to be searched. The user can then search for content matching the text indicated by the user in the video from the channel specified by the user and the time period specified by the user. This allows the user to quickly find at least one first target having the characteristics of the target to be found by the user, and display the first image of the found first target. Since the image search instruction is used to indicate a search for a first target in a certain first image, after receiving and responding to the image search instruction, it can be determined which of the displayed first images is to be searched for the first target, thereby accurately searching for all channels and times in which the target the user wants to detect is detected in the first video, without the user having to view the recorded video frame by frame, thereby improving the efficiency of target detection in the recorded video.

[0013] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0015] Figure 1 A first schematic diagram of the target detection method provided in an embodiment of the present application;

[0016] Figure 2 A first example diagram of an interface for displaying a first image provided in an embodiment of the present application;

[0017] Figure 3 A second schematic diagram of the target detection method provided in an embodiment of the present application;

[0018] Figure 4 This is a first example diagram of the emergency modeling interface provided in an embodiment of the present application;

[0019] Figure 5 A third schematic diagram of the target detection method provided in an embodiment of the present application;

[0020] Figure 6 This is a second example diagram of the emergency modeling interface provided in an embodiment of the present application;

[0021] Figure 7 This is a third example diagram of the emergency modeling interface provided in an embodiment of the present application;

[0022] Figure 8 This is a fourth example diagram of the emergency modeling interface provided in an embodiment of the present application;

[0023] Figure 9 A fourth schematic diagram of the target detection method provided in an embodiment of the present application;

[0024] Figure 10 This is an example diagram of the progress display interface provided in the embodiment of the present application;

[0025] Figure 11 A second example diagram of an interface for displaying a first image provided in an embodiment of the present application;

[0026] Figure 12 This is a third example diagram of an interface for displaying a first image provided in an embodiment of the present application;

[0027] Figure 13 This is a fourth example diagram of an interface for displaying a first image provided in an embodiment of the present application;

[0028] Figure 14aThis is a first example diagram of a text input interface provided in an embodiment of the present application;

[0029] Figure 14b This is a second example diagram of a text input interface provided in an embodiment of the present application;

[0030] Figure 15 This is a third example diagram of the text input interface provided in the embodiment of the present application;

[0031] Figure 16 An example diagram of the display interface provided in the embodiment of the present application;

[0032] Figure 17 An example diagram of a preset display window provided in an embodiment of the present application;

[0033] Figure 18 This is an example diagram of the update interface of the preset matching algorithm provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0035] First, the professional terms in the embodiments of this application are explained:

[0036] Image structuring: refers to the process of using computer vision and artificial intelligence technologies to convert unstructured image content (such as pixels and color blocks) into structured data with clear meaning that can be understood, queried, and analyzed by computers.

[0037] ITR (Image Text Retrieve) model: It is the core technology of cross-modal retrieval and can achieve bidirectional semantic alignment and matching between images and text;

[0038] Modeling: In this article, it refers to extracting target features detected in video images through a pre-trained model, and establishing model data for each target based on the extracted features. The modeling process can be regarded as the process of converting the features of the target into computer-calculated data. Moreover, in different scenarios, there may be different requirements for modeling due to different purposes or requirements of target detection, which in turn makes the data obtained by modeling different. For example: the characteristics of the target include one or more of the following features: multimodal features (features that integrate two or more different modal information such as images, text, audio, depth maps, etc.), unimodal features (color features, texture features, shape features, local features, depth features, etc.) or behavioral features, etc., and examples will not be given here one by one.

[0039] Model library: used to store model data related to target information generated after target modeling, such as target ID (identity document), target type, target coordinates, etc.

[0040] Vocabulary detection library: The device has a preset vocabulary library integrated inside to detect and filter some input content that does not meet user requirements and give user prompts.

[0041] In related technologies, detecting targets using video footage typically requires obtaining an image of the target in advance. A target detection algorithm is then used to extract the target's features. The target is then detected frame by frame within the video footage, and similarity is calculated. Only those candidates that meet the similarity criteria are identified as the target. For example, if a white pet dog is lost and its path needs to be determined within the video footage, an image of the dog is required before detection can be performed within the video footage.

[0042] If a photo of the pet dog cannot be provided, target detection cannot be performed directly according to the above method. The only way is to watch the video frame by frame to determine the movement path of the pet dog, and the detection efficiency is low.

[0043] In order to improve the efficiency of target detection in video recording, the first aspect of the embodiment of the present application provides a target detection method, such as Figure 1 The first schematic diagram of the target detection method provided in the embodiment of the present application is shown, and the method includes the following steps:

[0044] Step S10, receiving an input of a search condition, and displaying a first image of at least one first target hit by the search condition;

[0045] The search conditions include a first channel, a first time period, and a first text; the first image is from a first video; the first video is from a video of the first channel within a first time period; the first target is a target that matches the first text; and the first text includes: a subtext for indicating a target category and a subtext for indicating features of the target;

[0046] Step S20: receiving an image search instruction, the image search instruction being used to instruct a search for a first target in a first image, and, in response to the image search instruction, displaying target information of each second target detected in the first video;

[0047] The target information includes: a timestamp indicating the video frame in which the second target was detected and information about the channel from which it originated. The second target is a target that matches the first target indicated by the image search instruction in the first image indicated by the image search instruction. It is understood that, ignoring errors in the image search, the second target should be the same as the first target. For example, assuming that a first image is a photo of puppy A taken at 9:00 a.m. on January 1st, and the first target is puppy A, then when searching for this first target, all second targets obtained should be puppy A, except that the second targets also include puppy A appearing in images other than the first image.

[0048] According to the embodiment of the present application, text, channel, and time period are combined as search conditions. In this way, when searching for a target, the user only needs to enter text that matches the characteristics of the target to be searched, as well as the channel and time period to be searched. Then, the user can search for content that matches the text indicated by the user in the video of the time period specified by the user in the channel specified by the user. At least one first target having the characteristics of the target to be searched by the user can be quickly found, and the first image of the found first target can be displayed. Since the image search instruction is used to indicate the search for the first target in a certain first image, after receiving and responding to the image search instruction, it can be determined which first target in the displayed first image is to be searched, so that all channels and times where the target the user wants to detect is detected can be accurately searched in the first video without the user having to view the recorded video frame by frame, thereby improving the efficiency of target detection in the recorded video.

[0049] Still taking the target as a white puppy as an example, by applying the method provided in the embodiments of the present application, even if a photo of the white puppy cannot be provided, the characteristics of the target can be described by the first text. For example, if the first text is "white puppy", images with characteristics matching "white puppy" can be searched in the video recordings from the specified time period and the specified channel and the found images can be displayed to the user, so that the user can accurately select the target from the displayed images for further accurate matching and determine the time when the puppy appeared and the channel from which the video originated.

[0050] The above steps S10 and S20 are described in detail below:

[0051] It is understandable that the search condition is input by the user, and the first channel, first time period and first text included in the search condition are all specified by the user. Since a video processing device is usually connected to multiple image acquisition devices, each image acquisition device transmits the acquired image data to the video processing device through its corresponding channel. After receiving the image data transmitted by the image acquisition device through its respective channel, the video processing device stores the received image data in the video processing device. Therefore, the first video in the embodiment of the present application can refer to the video within the first time period received through the first channel among the image data historically stored by the video processing device. In this case, the first time period is a historical time period; it can also refer to the video within the first time period received by the video processing device through the first channel in real time. In this case, the first time period is the current time period. The current time period can refer to a time period starting at the current moment and ending at a certain moment in the future, or a time period starting at a certain moment in history and ending at the current moment.

[0052] The number of first channels included in the search condition can be one or more, and the number of first time periods can be one or more, depending on the number of first channels and the first time periods selected by the user. Each first time period can be one hour, one day, or other lengths.

[0053] The first text is used to describe the characteristics and category of the target. The characteristics of the target include but are not limited to appearance characteristics, movement characteristics, height, clothing color, clothing style, etc. If the category of the target is human, the characteristics of the target may also include facial features.

[0054] The first channel, the first time period, and the first text may be input in the same interface. For example, the user may select the first channel from the channel list displayed on the same interface, select the first time period from the time range displayed on the interface, and enter the first text in the input box displayed on the interface. In this example, the first channel, the first time period, and the first text are all input by the user in the same interface.

[0055] In another possible implementation, the first channel, first time period, and first text may be entered in different interfaces. For example, the user may select the first channel from the channel list displayed on the first display interface, select the first time period based on year, month, day, hour, and minute on the second display interface, and enter the first text on the third display interface.

[0056] It is understandable that the first target hit by the search condition refers to the first target that matches the first text in the video originating from the first channel in the first time period. There may be only one first target hit by the search condition, or there may be multiple first targets. In the case where there are multiple first targets hit by the search condition, when displaying the first image of the first target hit by the search condition, you can display only the first image of one first target hit by the search condition, or you can display the first images of a preset number of first targets hit by the search condition, or you can display the first images of all first targets hit by the search condition. For example, assuming that there are 100 first targets hit by the search condition, you can display the first images of these 100 first targets, or you can display the first images of only 10 of them, or you can display the first image of only 1 of them. The preset number is set by the user according to demand, and the embodiments of the present application do not limit this.

[0057] For the first target hit by the search condition, only a part of the first image of the first target may be displayed, or the entire first image of the first target may be displayed.

[0058] It can be understood that since the first video frame image is determined from the first video, there may be multiple video frames in which the first target is detected in the first video. In order to save computing resources, the video frame with the highest score among all video frames in which the same first target appears can be used as the first video frame image. For example, among all video frames in which the first target is detected in the first video, the video frame in which the first target has the highest clarity and is most forward of the image acquisition device can be used as the first video frame image.

[0059] See also Figure 2 , Figure 2The figure shows a first example of an interface for displaying a first image provided by an embodiment of the present application, in which multiple first images of first targets are displayed. In order to facilitate the user to determine whether the search result matches the first text, the content of the first text, such as "white puppy", is also displayed in the display interface. A first time period, such as "today", "three days", and "one week" is also displayed. The user can also select the first time period by himself through "customization", and the first time period can be a continuous time period or a time period composed of multiple discontinuous sub-time periods. For example, the first time period can be from 8:00 to 10:00 every morning from January 1 to January 3, and the user can also view and modify the first channel by clicking the displayed "channel" control.

[0060] It can be understood that the above step S10 is actually the process of inputting text to search images, and inputting text to search images is a cross-modal retrieval. In order to achieve cross-modal retrieval, a neural network model can be used to extract visual features from the image of each target detected, and establish a correspondence between each target and its corresponding visual features. A language processing model is used to extract key semantic information in the text, such as nouns: people, objects, animals, adjectives: attributes of objects (red), actions or spatial relationships: riding, wearing, etc. Finally, a multimodal pre-training model is used to learn how to map the targets in the image and the text describing each target into the same mathematical space. In this way, after the user enters the text, the target to which the visual feature with the highest similarity to the vector of the text content belongs will be used as the search result. A large model can also be used for cross-modal text-to-image search, such as an ITR model, or other large models for cross-modal search.

[0061] Since there may be multiple first objects in the first video that match the first text, some of them may not be the object the user is interested in. Through the display interface, the user can select the object of interest. That is, in step S20 above, the first object in the first image indicated by the image search instruction is the object of interest to the user.

[0062] An object that matches the first object in the first image indicated by the image search instruction is an object whose similarity to the first object in the first image indicated by the second search instruction is greater than a preset similarity threshold. The preset similarity threshold is set based on user needs and actual experience and is not limited in this embodiment of the present application.

[0063] When the second target matching the first target in the first image indicated by the image search instruction is determined, it is possible to determine which channel each second target originates from and at which time it was captured, that is, to determine the target information of each second target.

[0064] To improve search efficiency, target detection can be performed on the image data in advance, and a model of each detected target can be established. It is understandable that each time a first text is input for search, the entire first video must be processed first, target detection performed on each frame, and features of the detected targets extracted. The similarity between the first text and these temporarily extracted features is then calculated one by one, resulting in a large amount of computation and a long time consumption.

[0065] Based on this, see Figure 3 , Figure 3 This is a second schematic diagram of the target detection method provided in an embodiment of the present application, the method comprising the following steps:

[0066] Step S101, determining first targets hit by a search condition among targets pre-modeled from a first video;

[0067] Step S102, displaying a portion or all of the first images of each determined first target;

[0068] Step S20: receiving an image search instruction, the image search instruction being used to instruct a search for a first target in a first image, and, in response to the image search instruction, displaying target information of each second target detected in the first video;

[0069] Step S101 and step S102 are specific implementations of step S10.

[0070] In the above step S101, by modeling the first video in advance, a correspondence between each target detected in the first video and its characteristics can be established. After the user enters the search conditions, the similarity between the first text and the characteristics of each target is calculated, and each first target hit by the search conditions can be quickly determined, and part or all of the image of each first target can be displayed through S102.

[0071] In the above step S101, when determining the first targets hit by the search conditions among the targets obtained by pre-modeling the first video, the similarity between the targets obtained by modeling the first video and the first text included in the search conditions is compared, and the targets whose similarity with the first text is greater than a preset similarity threshold are determined as the first targets.

[0072] For example, assuming that the targets obtained by modeling the first video are target 1, target 2, target 3, and target 4, if the targets whose similarity with the first text is greater than the preset similarity threshold are target 1 and target 3, then the above step S101 determines that the first targets are target 1 and target 3 from "target 1, target 2, target 3, target 4".

[0073] In the above-mentioned step S102, part or all of the first images of each determined first target are displayed. It can be displaying part of the first image of each determined first target, or displaying all of the first images of each determined first target. It can also be displaying all of the first images of the first target whose similarity with the first text is greater than a preset similarity threshold, or displaying part of the first image of the first target whose similarity threshold with the first text is less than a preset similarity threshold. All of these are possible and are not limited to this in the embodiments of the present application.

[0074] It is understandable that when displaying the first images of the determined first targets in step S102, if the total number of first images of all determined first targets is small, the user may not be able to accurately find the target of interest. In this case, it is necessary to reanalyze the first video to determine whether there is a first target of interest to the user.

[0075] Based on this, Figure 3 As shown, the method further includes the following steps:

[0076] Step S30 , if the total number of first images of the first targets hit by the search condition is not greater than a preset number threshold, then displaying an emergency modeling control;

[0077] Step S40, in response to the interactive instruction of the emergency modeling control, modeling at least the unmodeled targets in the first video, and determining, from the newly modeled targets, first targets that are hit by the search condition as third targets, and displaying part or all of the first images of the third targets;

[0078] In the above step 30, the preset number threshold is set based on actual experience, for example, it can be 0, 1, or other natural numbers.

[0079] If the total number of the first images is not greater than the preset number threshold, it is considered that the number of the first images is small and there may not be a first target that the user is interested in, and the emergency modeling control is displayed to the user. Figure 4 The figure shows a first example of an emergency modeling interface provided by an embodiment of the present application. This interface displays emergency modeling controls. Furthermore, to facilitate user confirmation of the correctness of the first content and first time period entered, the interface also displays the user-determined first time period and displays the first content entered by the user in the search box.

[0080] After the user clicks the emergency modeling control, at least the unmodeled targets in the first video are modeled. Specifically, the targets in the first video may be re-detected and each detected target modeled separately. Alternatively, the targets in the first video may be re-detected, the unmodeled targets identified, and only those unmodeled targets are modeled. The first targets that are hit by the search criteria are then determined from the newly modeled targets as third targets, and part or all of the first image of each third target is displayed.

[0081] In the above step 40, part or all of the first images of each determined third target are displayed. It can be displaying part of the first image of each determined third target, or displaying all of the first images of each determined third target. It can also be displaying all of the first images of the third targets whose similarity with the first text is greater than a preset similarity threshold, or displaying part of the first images of the third targets whose similarity threshold with the first text is less than a preset similarity threshold. All of these are possible and are not limited to this in the embodiments of the present application.

[0082] It is understandable that in order to increase the number of first images of the first target hit by the search condition, the preset similarity threshold between the first text and the first target can be adjusted when determining each first target hit by the search condition from the new target obtained by modeling. Exemplarily, assuming that the preset similarity threshold between the first text and the first target is 95% initially, if the total number of first images of the first target determined from all targets according to the preset similarity threshold of 95% is less than the preset number threshold, an emergency modeling control is displayed to the user, and after the user interacts with the emergency modeling control, the target in the first video is re-modeled to obtain a new target, and the preset similarity threshold between the first text and the first target is set to 80%, and the target hit by the search condition is determined from the obtained new targets.

[0083] According to an embodiment of the present application, when the total number of first images of the first target is less than a preset threshold, the target in the first video is re-modeled through an emergency modeling control and the first target is determined based on the modeled target, thereby avoiding missed detection and improving the accuracy of target detection.

[0084] However, due to the limited computing resources of computers, if there is an ongoing emergency modeling task when the interactive instruction of the emergency control modeling component is received, it may be impossible to model the target in the first video. Based on this, in a possible implementation method, Figure 5 The third schematic diagram of the target detection method provided in the embodiment of the present application is shown, which includes the following steps:

[0085] Step S101, determining first targets hit by a search condition among targets pre-modeled from a first video;

[0086] Step S102, displaying a portion or all of the first image of each first target;

[0087] Step S20: receiving an image search instruction, the image search instruction being used to instruct a search for a first target in a first image, and, in response to the image search instruction, displaying target information of each second target detected in the first video;

[0088] Step S30 , if the total number of first images of the first targets hit by the search condition is not greater than a preset number threshold, then displaying an emergency modeling control;

[0089] Step S401: In response to an interactive instruction of the emergency modeling control, if there is no ongoing emergency modeling task, at least one target in the first video that has not yet been modeled is modeled, and first targets that are hit by the search condition are determined from the newly modeled targets as third targets, and part or all of the first image of each third target is displayed;

[0090] The first image of the third target can be displayed in addition to the first image of the first target, that is, the first images of the first and third targets can be displayed on the same interface. Alternatively, the first image of the third target can be displayed independently, for example, the first image of the third target can be displayed in a new interface, or the first image of the first target can be replaced with the first image of the third target.

[0091] Step S50, in response to the interactive instruction of the emergency modeling control, if there is an ongoing emergency modeling task, displaying the video modeling information of the video based on the ongoing emergency modeling task and the video modeling information of the first video;

[0092] Step S60, in response to the cancel modeling instruction of the displayed video modeling information, stopping the ongoing emergency modeling task and returning to step S401;

[0093] Step S70 , in response to the continue modeling instruction of the displayed video modeling information, continue the ongoing emergency modeling task.

[0094] The video modeling information includes one or more of the following information: estimated modeling completion time, the channel from which the video originates, the time period of the video, and the modeling completion progress;

[0095] For example, in Figure 4 In the interface shown, the user clicks the “Emergency Modeling Control” to start emergency modeling. During the emergency modeling process, the user can also view information such as modeling progress, such as Figure 6Shown is a second example diagram of the emergency modeling interface provided in an embodiment of the present application. The interface displays the estimated modeling completion time, channel, video time period and modeling completion progress of the video on which the ongoing emergency modeling task is based, as well as the estimated modeling completion time, channel and video time period of the first video.

[0096] It is understandable that Figure 6 This is only one possible example of an emergency modeling interface. In other possible embodiments, the emergency modeling interface may only include the estimated modeling completion time of the video on which the ongoing emergency modeling task is based and the estimated modeling completion time of the first video. It may also only include the time period of the video on which the ongoing emergency modeling task is based and the time period of the first video. It may also include one or more of the estimated modeling completion time, channel, and time period.

[0097] At this time, if the user wants to prioritize the new urgent modeling task, the user responds to the cancel modeling instruction of the displayed video modeling information, stops the ongoing urgent modeling task, and returns to execute the above step S401;

[0098] After completing the ongoing emergency modeling task, or after the user wants to stop the new emergency modeling task, the ongoing emergency modeling task can also be continued. Based on this, in response to the continue modeling instruction of the displayed video modeling information, the ongoing emergency modeling task is continued. For example, Figure 6 As shown, the user can click the "OK" control in the emergency modeling task window to perform a new emergency modeling task, or click the "Cancel" control in the emergency modeling task window to continue the ongoing emergency modeling task.

[0099] According to an embodiment of the present application, when a new emergency modeling instruction is received while there is an ongoing emergency modeling task, the video modeling information of the video on which the ongoing emergency modeling task is based and the video modeling information of the first video are displayed to the user, so that the user can select the emergency modeling task that needs to be executed first.

[0100] Since the process of emergency modeling of the first video may be long, it is necessary to wait until all targets are modeled in the first video and the first target matching the first text is determined before the first image of the first target can be displayed to the user, resulting in a poor user experience.

[0101] Based on this, in a possible implementation, in response to a continue search instruction during the emergency modeling of the first video, while continuing to model the targets that have not yet been modeled in the first video, the first target is determined from the targets that have been modeled from the first video, and the first image of each determined first target is displayed. The completion progress of the ongoing emergency modeling task, the estimated time to complete the modeling, and the "end" control for the user to instruct the end of the emergency modeling task and the "search" control for the user to instruct the determination of the first target from the targets that have been modeled from the first video can also be displayed. During the emergency modeling of the first video, the user can click the "search" control to determine the first target from the targets that have been modeled from the first video, and the user can also click the "end" control to end the emergency modeling of the first video.

[0102] For example, Figure 7 The third example diagram of the emergency modeling interface provided by the embodiment of the present application is shown. In the emergency modeling, it is assumed that 30% of the modeling of the first video has been completed and the estimated remaining time is 2 minutes. The user clicks the "Search" control, and the first target can be determined from the targets obtained by 30% modeling in the first video, and the target is selected as follows. Figure 8 The first image of the first target is shown. It can be understood that Figure 8 This is only an example. In actual display, the displayed first images are not completely the same.

[0103] Using the embodiments of the present application, while the first video is being urgently modeled, a first target is identified from the targets already modeled from the first video, and a first image of the first target is displayed. This allows the first image of the first target to be provided to the user in real time, reducing user waiting time and improving the user experience. Furthermore, after viewing a portion of the first image, the user may terminate unnecessary searches prematurely, thus saving computing resources.

[0104] Since different image acquisition devices transmit captured image data to the video processing device through different channels, and different image acquisition devices are usually set at different locations, the location where the user appears is reflected in the image data as the channel of the video. For example, assuming that the channel through which image acquisition device 1 transmits image data to the video processing device is channel 1, if the target of interest to the user appears at position A during time period 1, and image acquisition device 1 is used to capture the image of position 1, the user needs to view the video received by the video processing device through channel 1 during time period 1. It will be understood that the video of the first channel in the first time period in this application refers to the video captured by the image acquisition device that transmits image data through the first channel during the first time period.

[0105] In order to facilitate users to understand the channel and time period of each video, in a possible implementation, as shown in FIG. Figure 9 FIG. 3 is a third schematic diagram of a target detection method provided in an embodiment of the present application, the method comprising the following steps:

[0106] Step S101, determining first targets hit by a search condition among targets pre-modeled from a first video;

[0107] Step S102, displaying a portion or all of the first image of each first target;

[0108] Step S20: receiving an image search instruction, the image search instruction being used to instruct a search for a first target in a first image, and, in response to the image search instruction, displaying target information of each second target detected in the first video;

[0109] Step S30 , if the total number of first images of the first targets hit by the search condition is not greater than a preset number threshold, then displaying an emergency modeling control;

[0110] Step S401: In response to an interactive instruction of the emergency modeling control, if there is no ongoing emergency modeling task, at least one target in the first video that has not yet been modeled is modeled, and first targets that are hit by the search condition are determined from the newly modeled targets as third targets, and part or all of the first image of each third target is displayed;

[0111] Step S50, in response to the interactive instruction of the emergency modeling control, if there is an ongoing emergency modeling task, displaying the video modeling information of the video based on the ongoing emergency modeling task and the video modeling information of the first video;

[0112] Step S60, in response to the cancel modeling instruction of the displayed video modeling information, stopping the ongoing emergency modeling task and returning to step S401;

[0113] Step S70 , in response to the continue modeling instruction of the displayed video modeling information, continue the ongoing emergency modeling task.

[0114] Step S801: In response to a progress display instruction, a channel list including various channels is displayed, and a channel of a video being urgently modeled is identified in the channel list;

[0115] Step S802: In response to a channel list selection instruction, display a time period list, which includes the shooting time period of each video of the channel indicated by the selection instruction;

[0116] Step S803: If the channel indicated by the selection instruction is the channel of the video being urgently modeled, the shooting time period of the video being urgently modeled is marked in the time period list.

[0117] The above steps S101 to S70 are described above and will not be repeated here.

[0118] The following is combined with Figure 10 The above steps S801 to S803 are explained. Figure 10 Shown is an example diagram of a progress display interface provided in an embodiment of the present application.

[0119] In response to the progress display instruction, the display Figure 10 The interface shown in the figure displays a channel list, which displays each channel. The user can select the channel they want to view from the displayed channels. After the user selects the channel they want to view, the user is presented with a time period list, which displays the time period of the video of the selected channel.

[0120] In order to help users understand the channel and time period of the video on which the ongoing emergency modeling task is based, the channel of the video on which the emergency modeling task is being performed will be identified in the channel list. When the channel selected by the user is the channel of the video on which the ongoing emergency modeling task is based, the time period of the video on which the ongoing emergency modeling task is based will also be identified in the time period list. For example, Figure 10 As shown, the channels marked with “*” in the channel list are the channels of the videos on which the ongoing emergency modeling task is based, and the time periods marked with “#” in the time period list are the time periods of the videos on which the ongoing emergency modeling task is based.

[0121] It can be understood that when identifying the channel of the video on which the ongoing emergency modeling task is based in the channel list and identifying the time period of the video on which the ongoing emergency modeling task is based in the time period list, not only "*" and "#" can be used for identification, but other styles of identification can also be used, such as using different colors for identification.

[0122] In order to facilitate users to view the analysis progress of videos in each channel and each time period, in a possible implementation, when displaying the time period list to the user, the analysis progress of the videos shot in each shooting period will also be displayed in the time period list. For example, Figure 10 As shown, the number "56" displayed on the time period indicates that the emergency modeling completion progress of the video in this time period is 56%.

[0123] By using the embodiment of the present application, by displaying a time period list and a channel list to the user, the user can select the channel and time period to be viewed according to needs. Moreover, by identifying the channel and time period of the video on which the ongoing emergency modeling task is based, the user can intuitively understand the channel and shooting time period of the video on which the ongoing emergency modeling task is based, thereby further improving the user experience.

[0124] It is understandable that the same target may be captured by the same image capture device at different time periods. Consequently, repeated first targets may appear in the displayed first images, forcing users to identify the target of interest amidst a large amount of duplicate content, impacting the user experience. To reduce the impact of redundant information on users, in one possible implementation, if the number of first images of the same target exceeds a preset threshold, only some of the first images of that target are displayed.

[0125] When displaying part of the first image of each first target, one first image may be displayed for each first target, two first images may be displayed for each first target, or a preset number of first images may be displayed for each first target.

[0126] In a possible implementation, when displaying the first image of at least one first target hit by the search condition, the first image can be displayed directly on the Figure 2 The interface shown displays at least one, and / or at most N, first images of each first target. N is a preset positive integer, and the specific value can be determined empirically. In one possible embodiment, to minimize redundant information interference, N is set to 1, meaning that only one first image of each first target is displayed.

[0127] In another possible embodiment, when displaying the first image of at least one first target hit by the search condition, the respective file covers of some or all of the first targets hit by the search condition are displayed. The file to which the file cover belongs can be represented in the form of a folder or in other ways. For example, Figure 11 The second example diagram of an interface for displaying a first image provided by an embodiment of the present application uses different profile covers to display first images of different first targets. Similarly, to facilitate the user's determination of whether the search result matches the first text, the display interface also displays the content of the first text, such as "white puppy," and a first time period, such as "today," "three days," or "one week." The user can also select the first time period by clicking "Customize," and can also view and modify the first channel by clicking the displayed "Channel" control.

[0128] Since all the first images of the first target are not displayed at this time, in order to facilitate the user to view the complete information, after receiving the expansion instruction, in response to the expansion instruction, some or all images included in the file represented by the file cover indicated by the expansion instruction can be displayed. Figure 12 The third example diagram of the interface for displaying the first image provided by the embodiment of the present application is shown. Figure 11 Select the folder of the first target 1 in the interface shown, then Figure 12 The interface shown displays part or all of the first image of the first target 1. Furthermore, the interface may also display the content of the first text, such as "white puppy," and a first time period, such as "today," "three days," or "one week." Users may also select the first time period through "Customize," and may view and modify the first channel by clicking the displayed "Channel" control. To help users understand whether the target in the displayed first image is the target they are interested in, the display interface also displays the degree of match between the target in the first image and the first text. For example, "89%" in the figure indicates that the target in the first image matches the first text "white puppy" by 89%.

[0129] By adopting the embodiment of the present application, by displaying the file covers of the respective first targets and then expanding to display the first images of the first targets that the user is interested in when the user needs them, the interference of redundant information on the user can be reduced and the user experience can be improved.

[0130] In order to help users make decisions, when displaying the first image of at least one first target hit by the search conditions, the image information of each first image can also be displayed at the same time. The image information includes one or more of the shooting time, channel, and the similarity between the first target in each first image and the first text.

[0131] It is understandable that when displaying the image information of each first image, the image information of each first image is displayed at the corresponding position of each first image. Figure 13 The fourth example diagram of the interface for displaying the first image provided by the embodiment of the present application is shown. In each displayed first image, the time when the first image was captured, the channel from which the first image originated, and the degree of matching between the first target in the first image and the first text are respectively displayed. Figure 13 “2025.1.1 12:10 Channel 1” displayed on the first image indicates that the first image was taken by the image acquisition device corresponding to Channel 1 on January 1, 2025, and “89” indicates that the matching degree between the first target appearing in the first image and the first text “white puppy” is 89%.

[0132] The degree of matching between the first object in the first image and the first text refers to the similarity between the image features of the first object in the first image and the text features of the first text. The similarity between the text features and the image features can be calculated using a cross-modal similarity calculation method, such as an ITR model or other cross-modal similarity calculation methods.

[0133] By adopting the embodiment of the present application, by displaying the image information of each first image, it is convenient for users to select targets of interest based on time, channel or matching degree, helping users make decisions, and thus improving user experience.

[0134] exist Figure 13 When displaying each first image in the interface shown, the user can choose to display them in order of high to low or low to high similarity between the first target in each first image and the text features of the first text, or can choose to display them in chronological order or in reverse chronological order.

[0135] However, since the number of first images is relatively large, in order to avoid users from manually flipping through the massive first images and to improve the efficiency of users in determining the target they want to search for, after each first image is displayed, the user can customize the minimum similarity between the target of the first image to be displayed and the text features of the first text. Figure 13 As shown, the user can set the similarity by dragging the similarity adjustment control. For example, if the user drags the similarity adjustment control to a similarity of 60, then Figure 13 The minimum similarity between the targets of all first images displayed in the interface and the text features of the first text is 60%. The user can also directly enter the minimum similarity they want to set in the interface and adjust the currently set minimum similarity through the increase and decrease symbols or adjustment arrows.

[0136] It is understandable that, due to the different categories of targets, the image features of the targets in the image are different. When describing targets of different categories, the features contained in the text are also different. For example, the text describing a puppy usually includes features such as the puppy's coat color, limbs, tail, and ears, but the text describing a vehicle usually includes features such as the vehicle's color, model, and logos on the vehicle. In order to make the target detection results more accurate, the target types in the training data used when detecting models of different categories of targets are also different. For example, the model used for vehicle detection is trained using a large number of pictures containing vehicle data, while the model used for bicycle detection is trained using a large number of pictures containing bicycles.

[0137] Therefore, in order to improve the accuracy of target detection, in a possible implementation, the category of the target may be determined when the first text is input.

[0138] In one possible implementation, the target detection method provided in this application includes the following steps:

[0139] Step S1001: Display multiple category cards representing different categories in the vicinity of the search box;

[0140] Step S1002, in response to the target category selection instruction, determining the category card selected by the target category selection instruction, taking the category represented by the selected category card as the first category, and displaying the respective tabs of the characteristics of the targets in the first category;

[0141] Among them, each feature's tab page displays its own sub-features;

[0142] Step S1003, in response to the second selection instruction of the displayed sub-feature, generating a first text including a first sub-text and a second sub-text;

[0143] The first subtext is used to describe all sub-features indicated by the second selection instruction, and the second subtext is used to describe the first category.

[0144] Step S10, receiving an input of a search condition, and displaying a first image of at least one first target hit by the search condition;

[0145] Step S20 , receiving an image search instruction, where the image search instruction is used to instruct to search for a first target in a first image, and, in response to the image search instruction, displaying target information of each second target detected in the first video.

[0146] The above steps 10 and S20 are described above and will not be repeated here.

[0147] In the above step S1001, the search box is used for the user to input the first text, and a plurality of category cards representing different categories are displayed in the vicinity of the search box. Figure 14a , Figure 14a The first example diagram of the text input interface provided by the embodiment of the present application is shown. In this interface, multiple category cards are displayed, such as "Find People", "Find Animals", "Find Cars", "Find Bicycles", and "Find Items". After the user selects a category, the possible characteristics of the target in that category will be displayed. For example, if the target category is "People", the possible characteristics displayed are "wearing a hat, wearing glasses, wearing a mask, carrying a bag, and backpack". The user can select the corresponding characteristics according to actual needs.

[0148] Different categories are pre-set, such as people, animals, and objects. The characteristics of the first category of targets are pre-set, and the characteristics of targets in different categories are different. After the user selects the first category, the respective tabs of the characteristics of the first category of targets will be displayed so that the user can select different sub-features in different tabs. For example, Figure 14b The figure shows the second example of the text input interface provided in an embodiment of the present application. When the user selects the category card as "Find People", the respective tabs of each feature of the person are displayed below the search box, such as "Tops", "Bottoms", "Accessories", "Behaviors", etc., so that the user can select from the sub-features displayed in different tabs.

[0149] Because the sub-features selected by the user are discrete and not a complete text, to facilitate computer understanding of the user's selected features, the user's selected features can be converted into text containing each feature and the selected category in step S1003. For example, assuming the user selected the category "people," and the sub-features displayed in the accessories tab are selected as "backpack," "wearing a mask," and "wearing a hat," the first text generated would be "people wearing a backpack, a mask, and a hat."

[0150] By using the embodiment of the present application, the user can generate a first text by selecting a category and features of the target, so that the most appropriate model can be selected for search, thereby improving the accuracy of target detection.

[0151] In order to facilitate users to flexibly change the characteristics of the target when entering the first text, after the user enters the first text through the search box, the first text is displayed in the search box, and the marks of each characteristic sub-text in the first text are displayed. The user can flexibly edit the first text in the search box.

[0152] Based on this, in one possible implementation, the target detection method provided in the embodiments of the present application includes the following steps:

[0153] Step S1001: Display multiple category cards representing different categories in the vicinity of the search box;

[0154] Step S1002, in response to the target category selection instruction, determining the category card selected by the target category selection instruction, taking the category represented by the selected category card as the first category, and displaying the respective tabs of the characteristics of the targets in the first category;

[0155] Step S1003, in response to the second selection instruction of the displayed sub-feature, generating a first text including a first sub-text and a second sub-text;

[0156] Step S1004: display the first text in the search box, and display a mark and a modification control on the first text;

[0157] The tag is used to mark each feature subtext. The feature subtext is a subtext in the first text used to describe a subfeature. Different feature subtexts are used to describe different subfeatures.

[0158] Step S1005 , in response to the third selection instruction, determining the feature subtext corresponding to the modification control indicated by the third selection instruction as the first feature subtext, and displaying a subfeature display interface;

[0159] The sub-feature display interface includes other sub-features on the same tab page as the sub-feature described by the first feature sub-text;

[0160] Step S1006 , in response to the displayed fourth selection instruction of other sub-features, replacing the first feature sub-text in the first text with the second feature sub-text to obtain a new first text;

[0161] The second feature subtext is a subtext used to describe other subfeatures indicated by the fourth selection instruction;

[0162] Step S10, receiving an input of a search condition, and displaying a first image of at least one first target hit by the search condition;

[0163] Step S20 , receiving an image search instruction, where the image search instruction is used to instruct to search for a first target in a first image, and, in response to the image search instruction, displaying target information of each second target detected in the first video.

[0164] The above steps 10 and S20, and steps S1001 to S1003 are described above and will not be repeated here.

[0165] In the above step S1004, when the first text is displayed in the search box, each characteristic sub-text and the corresponding modification control of each characteristic sub-text are marked in the first text. Figure 15 The third example diagram of the text input interface provided by the embodiment of the present application is shown. Assuming that the first text is "a person wearing a red shirt and a hat", when displaying the first text, the characteristic sub-texts "red shirt" and "wearing a hat" are marked in the first text.

[0166] The user can select and replace the marked feature sub-text through the above steps S1005 and S1006. For example, assuming that the feature sub-text corresponding to the modification control indicated by the third selection instruction is "red shirt", other sub-features on the same tab page as "red shirt" will be displayed, such as green, black, etc. After selecting "green", the "red shirt" in the first text is replaced with "green shirt", and a new first text is obtained, namely "a person wearing a green shirt and a hat".

[0167] By adopting the embodiment of the present application, the user can flexibly select each characteristic sub-text in the first text without having to re-enter a new first text, thereby improving the user experience.

[0168] Since the sub-features displayed in the tab page are limited, if the feature of the target that the user is interested in is not displayed in the tab page, the user can also directly enter a new first text through the search box without generating a new first text through the third selection instruction.

[0169] To increase user stickiness, each time a user enters a first text, the input can be recorded, and the user can be shown the first texts entered recently the next time the user enters a first text. In addition, when the user is unsure how to accurately describe the first text, the user can be recommended a first text obtained according to a preset recommendation algorithm.

[0170] Based on this, in a possible implementation, at least one candidate text will be displayed in the vicinity of the search box; wherein the candidate text includes text that has served as the first text within a historical period, and / or text determined according to a preset recommendation algorithm; the user can select the candidate text as the first text.

[0171] If there are a large number of texts that have served as the first text in the historical period and cannot be displayed in this interface, a page-turning control can be set in the interface, and the user clicks the page-turning control to display more first texts; or a preset number of first texts can be selected from the texts that have served as the first text in the historical period for display. Both are acceptable.

[0172] By using the embodiment of the present application, by displaying to the user the text or recommended text that has served as the first text in a historical period, the user can directly select the first text from the historical record or recommended text according to needs without having to re-enter the first text, thereby improving the user experience.

[0173] It is understood that the user may have pre-set a vocabulary detection library, and the first text input by the user needs to be matched with the vocabulary in the vocabulary detection library to determine whether the first text input by the user contains any of the vocabulary in the vocabulary detection library. The vocabulary in the vocabulary detection library is set by the user according to their needs, and the vocabulary detection library includes a plurality of pre-set words.

[0174] If the first text contains words from the vocabulary detection library, then the step of displaying the first image of at least one first target matched by the search criteria in step S10 cannot be performed. Instead, the user should be prompted to change the first text. For example, assuming "AA" is a word from the vocabulary detection library and the first text is "AABBCCDD", the user may be prompted with the message "Current content is not supported. Please try another description."

[0175] It is understandable that when displaying the first image of at least one first target hit by the search conditions, if the user needs to change the first text and search again, in order to improve the user experience, the first text is also displayed in the interface, and the user can directly edit the first text in the interface.

[0176] In this case, in response to the displayed first text edit instruction, the respective tabs for each feature of the first target category are displayed, with each tab displaying its own sub-features. In response to the displayed fifth selection instruction, a second text including the third and fourth sub-texts is generated as the new first text, and a first target matching the new first text is determined, and then the process returns to step S10. This allows the user to edit the first text directly in the interface displaying the first image, improving the user's interactive experience.

[0177] It can be understood that in the present application, when displaying the first image of each first target, the first video frame image in which the first target is detected in the first video can be displayed, or a sub-image of the area where the first target is located in the first video frame image can be displayed. This is both possible and the embodiments of the present application are not limited to this.

[0178] Since the image size of the sub-image of the area where the first target is located in the first video frame image is smaller than the size of the first video frame, in order to display more first preview images in the display interface and facilitate users to quickly determine the target of interest, when displaying the first image of each first target appearing in the first video, only the sub-image of the area where the first target is located in the first video frame image can be displayed, where the first video frame is the video frame in which the first target appears in the first video.

[0179] In this case, when the user selects a first image, the first video frame to which the first image belongs can be fully displayed.

[0180] Based on this, in a possible implementation, the method further includes:

[0181] In response to a zoom instruction for a displayed first image, a second image corresponding to the first image indicated by the zoom instruction is displayed; wherein the second image is a sub-image of a second region of the first video frame where the first object is located, and the first region is a proper subset of the second region. It is understood that the size of the second region can be smaller than or equal to the size of the first video frame. When the size of the second region is equal to the size of the first video frame, the second image is the first video frame itself.

[0182] The zoom instruction can be triggered by the user by single-clicking, double-clicking, or hovering after selecting the first image, or it can be triggered by the user by using a shortcut key after selecting the first image. This embodiment of the present application is not limited to this.

[0183] When displaying the second image corresponding to the first image indicated by the zoom instruction, it can be displayed in a new interface, or displayed on the upper layer of the first image indicated by the zoom instruction, or displayed in another location in the current interface. All of these are possible.

[0184] By using the embodiment of the present application, more first images can be displayed in the display interface by displaying a sub-image of the area where the first target is located in the first video frame image, which facilitates users to quickly identify targets of interest. Moreover, compared with enlarging and displaying the second image of each first image, enlarging and displaying the first image in response to a zoom instruction can greatly reduce the bandwidth consumption when loading each first image.

[0185] In this case, in order to facilitate the user to quickly identify the first target in the first image, in a possible implementation, while displaying the second image corresponding to the first image indicated by the zoom instruction, a target frame of the first target is displayed in the second image, where the first target is the target that matches the first text. For example, Figure 16 The figure shows an example of a display interface provided by an embodiment of the present application. When the second image 1 corresponding to the first image is magnified and displayed, the first target in the second image 1 is marked with a target frame. Figure 16 The dotted box in the second image 1 is the target box.

[0186] In order to further enhance the user experience, a video containing the first image may be played to the user when the user has a need.

[0187] Based on this, in a possible implementation, it may be that in response to the first playback instruction of the displayed first image, the first image indicated by the first playback instruction is used as the target image, the first video frame where the target image is located is determined as the target video frame, the sub-video in the first video that contains the target video frame and continuously detects the first target in the target image is determined as the first target video, and the first target video is displayed in a preset display window.

[0188] In another possible embodiment, in response to the second playback instruction of the displayed first image, the first image indicated by the second playback instruction is determined as the target image, the first video frame where the target image is located is determined as the target video frame, each first video frame in the first video that is continuous with the target video frame and detects the first target in the target image is obtained, each first image in each first video frame is obtained, the second target video is obtained, and the second target video is played in a floating window at the position where the target image is displayed.

[0189] In the above embodiment, the first play instruction and the second play instruction can be triggered by the user selecting the first image by single-clicking, double-clicking, or hovering, or by the user selecting the first image by using a shortcut key, and the present embodiment is not limited to this. It is understood that the triggering method of the first play instruction is different from the triggering method of the second play instruction, and the triggering method of the first play instruction and the triggering method of the second play instruction are different from the triggering method of the zoom instruction described above.

[0190] A sub-video in the first video that contains a target video frame and continuously detects the first target in the target image refers to a video composed of all adjacent video frames in the first video in which the first target in the target image is detected. For example, assuming that the first video includes 1000 video frames, for the first target 1, the first video frame corresponding to the target image is the 233rd video frame in the first video, and the first target 1 is detected in video frames 233 to 256 of the first video, then the first target video is the video composed of video frames 255 to 256.

[0191] The preset display window can be set in the interface for displaying each first image, or can be set in other display interfaces. For example, the preset display window can be set in the interface for displaying each first image. Figure 16 As shown, after the user double-clicks to select the first image 2, the corresponding first target video is displayed in the display window 3 on the right.

[0192] The first video frames in the first video that are continuous with the target video frame and detect the first target in the target image refer to the video frames in the first video that are adjacent to each other and detect the first target in all the video frames. For example, assuming that the first video includes 1000 video frames, for the first target 1, the first video frame corresponding to the target image is the 233rd video frame in the first video, and the first target 1 is detected in the 233rd to 256th video frames in the first video, then the first video frames are the 255th to 256th video frames. Then, the first images in the first video frames are played in the order of the first video frames to obtain the second target video, and the second target video is played in a floating window at the position where the target image is displayed.

[0193] The following takes the example of triggering the second play instruction in a hovering manner to illustrate the generation process of the second play instruction.

[0194] Step 1: upon detecting a hover event of a first image, displaying a playback control over the first image;

[0195] Step 2: Generate a second play instruction in response to an operation on the play control.

[0196] In the above step 1, the hover event for any displayed first image refers to the user moving the mouse pointer to any first image and briefly stopping there.

[0197] For example, assuming that the user moves the mouse pointer to the first image 2 and stays there for 1 second, a play control is displayed superimposed on the first image 2. When the user clicks or double-clicks the displayed play control, a second play instruction for the first image 2 is generated.

[0198] By using the embodiment of the present application, the first image to be played is selected by hovering, and the playback control of the first image will only be displayed when there is a need to play a certain first image. Compared with displaying a playback control for each first image, the number of first images that can be displayed in the interface can be increased, and the interface can be kept extremely simple, thereby improving the user's visual experience.

[0199] When the first target video is displayed in the preset window, since there may be video frames in which multiple targets are detected simultaneously in the first target video, in order to facilitate the user to accurately and quickly determine the first target when viewing the first target video, the first target can be marked with a target box in each video frame of the first target video.

[0200] It is understandable that if the user enters an incorrect first text, resulting in the first target in the displayed first image not being the target of interest to the user, but the user sees the target of interest when viewing the first target video, the user can re-enter the correct first text. In order to improve the detection speed, the target search function can also be turned on. In response to the target search function enable instruction, the target frame of each detected target is displayed in the video frame currently displayed and played in the preset display window. In response to the target selection instruction, each target that matches the target in the target frame indicated by the target selection instruction is detected in the first video as each fourth target, and the target information of each fourth target is displayed respectively. Among them, the target search function is a function of realizing target detection based on the target detection method provided in this application.

[0201] For example, Figure 17 The figure shows an example of a preset display window provided in an embodiment of the present application. When the first target video is displayed in the preset window 3, the target frame of the target detected in the video frame is also displayed in the currently displayed video frame. Figure 17 The dotted box in the preset window 3 is the target box.

[0202] Through the embodiments of the present application, users can change the detection target at any time according to their needs while viewing the first target video without having to re-enter the first text for retrieval, which improves the speed of target detection and further enhances the user experience.

[0203] It is understandable that if the features represented by the first text in the aforementioned step S10 are relatively broad, the number of displayed first images may be relatively large, resulting in the user being unable to quickly lock onto the target of interest.

[0204] In order to facilitate users to quickly lock on to the target they want to find, when displaying the first image, the controls corresponding to the features of the first target in the first image will also be displayed. In response to the control selection instruction of the displayed control, the image with the features corresponding to the control displayed by the control selection instruction is searched in the determined first image and displayed.

[0205] Displaying the features of the first target may include displaying all features of the first target, or displaying at least one feature that appears most frequently in the first image. For example, if more than 80% of the targets in the displayed first images are wearing hats, then the feature "wearing a hat" is displayed.

[0206] By using the embodiments of the present application, users can further search for targets of interest through the displayed features, thereby speeding up the process of locking onto targets of interest.

[0207] It is understandable that the aforementioned step 20 may be performed simultaneously with the steps of displaying the control corresponding to the feature possessed by the first target in the first image, searching for an image having the feature corresponding to the control displayed by the control selection instruction in the determined first image in response to the control selection instruction of the displayed control, and displaying the same, or they may be performed sequentially. In one possible embodiment, after performing the aforementioned step S10, the steps of displaying the control corresponding to the feature possessed by the first target in the first image, searching for an image having the feature corresponding to the control displayed by the control selection instruction in the determined first image in response to the control selection instruction of the displayed control, and displaying the same, may then be performed, and then step S20 may be performed.

[0208] Using the embodiment of the present application, two searches are first performed based on the text, so that the target in the displayed first image matches the target of interest to the user as much as possible, which facilitates the user to accurately select the target to be detected from the displayed first image, thereby accurately determining all channels and times where the target the user wants to detect appears in the first video, thereby improving the efficiency of target detection in the recorded video.

[0209] As can be seen from the above, the first target is a target whose similarity with the first text is greater than the preset similarity threshold. When matching the first target whose similarity with the first text is greater than the preset similarity threshold, matching is performed using a preset matching algorithm.

[0210] If the first target obtained according to the preset matching algorithm does not meet expectations, it may be that the parameters of the preset matching algorithm are not properly set. Based on this, in one possible implementation, in response to the labeling instruction, the first image indicated by the labeling instruction is determined as the labeled image, and the true value labeled for the labeled image is determined; then the preset matching algorithm is trained using the labeled image and the true value labeled for the labeled image, and the trained preset matching algorithm is used to re-match and obtain a new first target, and the process returns to execute the aforementioned step S10.

[0211] For example, Figure 18The figure shows an example of an update interface for a preset matching algorithm provided in an embodiment of the present application. For each displayed first image, the user can select "correct" or "wrong" to mark whether each displayed first image matches the first text. After marking, click the "retrain" control to train the preset matching algorithm using the marked image. In addition, the interface also displays the text content of the first text, such as "white puppy", and displays the query time as "2020 / 01 / 11 12:00:00-2020 / 01 / 12 13:00:00" and the query range as "Channel 1", indicating that an image of a white puppy was detected in the video shot by the image acquisition device of channel 1 during the period of 2020 / 01 / 11 12:00:00-2020 / 01 / 12 13:00:00.

[0212] By adopting the embodiment of the present application, by annotating the true value of the first image and using the annotated true value and the annotated image to train the preset matching algorithm, the accuracy of matching the first target can be improved, thereby improving the accuracy of target detection.

[0213] Furthermore, the updated preset matching algorithm may be used to re-match the first target, and the process returns to step S10 to make the displayed first image of the first target more closely match the first text, thereby improving the accuracy of target detection.

[0214] Based on the above, we can see that in this application, all targets detected in the first video will be modeled in advance. During the modeling process, if the target is an object, the name of the object will be marked. In this way, the user can enter the name of the object, display the image of the detected object, and select the object in the image. For example, if the user enters the name of the object as "book", all images with the detected "book" will be displayed, and the book will be selected in the image.

[0215] It's understandable that in related art, when searching for an item, video captures are taken at a user-set capture interval, and then a model is created for the target captured. When searching for an item, the user matches the modeling information with the item name to obtain and display the target image. If 100 captures show the same stationary item, the text search results will display 100 images, resulting in excessive redundant information presented to the user, impacting the user experience.

[0216] Based on this, in a possible implementation, the user can choose whether to turn on the static target filtering function. If turned on, in response to the static target filtering instruction, the determined first image is filtered according to the filtering rules indicated by the static filtering instruction, and the preset number of first images remaining after filtering are displayed.

[0217] In this case, the first image of at least one first target hit by the search condition is displayed in the aforementioned step S10, including: filtering out all images hit by the preset filter condition from all first images of at least one static first target hit by the search condition, as the filtered first images of the static first target; at least displaying all filtered first images of the static first target.

[0218] Through the embodiments of the present application, the first image of the static first target can be filtered when searching for an item, thereby reducing redundant information displayed to the user and improving the user experience.

[0219] It is understandable that, during the process of modeling the first video, if the same stationary object appears in multiple screenshots, it is impossible to determine which screenshot of the object to display during the object detection process after modeling. It may even be necessary to sequentially calculate the similarity between the object in each screenshot and the object of interest to the user, so as to display the screenshot with the most similar object, which consumes a lot of computing resources. Based on this, in one possible implementation, the process of pre-modeling the object in the first video in the aforementioned step S101 includes the following:

[0220] Step A: extracting video frames from the first video according to a preset frame extraction interval, and modeling static objects appearing in the extracted video frames;

[0221] Step B: determining a video frame in which a dynamic target appears from the first video, and modeling the dynamic target appearing in the determined video frame.

[0222] By extracting video frames according to a preset frame extraction interval to model static targets and model dynamic targets, the computing resources for static target matching can be reduced.

[0223] As can be seen from the above, the target detection method provided in the embodiment of the present application supports searching for images using text (hereinafter referred to as text search), searching for images using images (hereinafter referred to as image search), and also supports searching for objects. Text search, image search, static filtering, object search, etc. can all be regarded as different functions. In order to facilitate users to quickly open the functions of interest, in a possible implementation method, a function search interface is also displayed. In this interface, users can jump directly to the text search interface by entering the function name, such as "text search".

[0224] A second aspect of the embodiments of the present application provides a target detection method, the method comprising:

[0225] Receive input of search conditions and display a first image of at least one first target hit by the search conditions; wherein the search conditions include a first channel, a first time period and a first text, the first image comes from a first video, the first video comes from a video of the first channel within the first time period, the first target is a target that matches the first text, and the first text includes: a subtext for representing a target category and a subtext for representing characteristics possessed by the target.

[0226] Corresponding to the first aspect, a third aspect of the embodiments of the present application provides a video processing device, which may be a video recording device or other device for video processing. The video processing device is used to implement the target detection method described in the first and second aspects.

[0227] It should be noted that in the technical solution of this application, the operations involved in obtaining, storing, using, processing, transmitting, providing and disclosing user personal information are all carried out with the user's authorization.

[0228] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0229] Each embodiment in this specification is described in a related manner. Similar portions between the embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. In particular, the video processing device embodiment is generally similar to the method embodiment, so its description is relatively simple. For related portions, reference can be made to the description of the method embodiment.

[0230] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.

Claims

1. A target detection method, characterized in that: The method comprises: Receive input of search conditions and display a first image of at least one first target matched by the search conditions; wherein the search conditions include a first channel, a first time period, and a first text; the first image is from a first video; the first video is from a video of the first channel within the first time period; the first target is a target that matches the first text; and the first text includes: a subtext for indicating a target category and a subtext for indicating features possessed by the target; An image search instruction is received, where the image search instruction is used to instruct a search for a first target in a first image, and, in response to the image search instruction, target information of each second target detected in the first video is displayed respectively, wherein the target information includes: a timestamp of a video frame indicating that the second target is detected and information of a channel from which the video frame originates, the second target being a target that matches the first target in the first image indicated by the image search instruction.

2. The method according to claim 1, characterized in that The displaying of a first image of at least one first target hit by the search condition includes: Determining first targets hit by the search condition among targets obtained by pre-modeling the first video; displaying a portion or all of the first images of each of the determined first targets; The method further comprises: If the total number of first images of the first target hit by the search condition is not greater than a preset number threshold, displaying an emergency modeling control; In response to the interactive instructions of the emergency modeling control, at least the targets that have not been modeled in the first video are modeled, and the first targets hit by the search conditions are determined from the new targets obtained by modeling as the third targets, and part or all of the first images of the third targets are displayed.

3. The method according to claim 2, characterized in that The step of modeling at least an unmodeled target in the first video in response to the interactive instruction of the emergency modeling control comprises: In response to an interactive instruction of an emergency modeling control, if there is no ongoing emergency modeling task, modeling at least an unmodeled target in the first video; wherein the emergency modeling task is a modeling task performed in response to an interactive instruction of the emergency modeling control; or The method further comprises: In response to an interactive instruction of the emergency modeling control, if there is an ongoing emergency modeling task, displaying video modeling status information of the video on which the ongoing emergency modeling task is based and video modeling status information of the first video, wherein the video modeling status information includes one or more of the following information: an estimated modeling completion time, a channel from which the based video originates, a time period of the based video, and modeling completion progress; In response to a cancel modeling instruction in the displayed video modeling status information, the ongoing emergency modeling task is stopped, and the steps of modeling at least the unmodeled targets in the first video, determining the first targets hit by the search condition from the new targets obtained by the modeling as the third targets, and displaying at least part or all of the first images of the third targets; and / or, In response to the continue modeling instruction of the displayed video modeling information, the ongoing emergency modeling task is continued.

4. The method according to claim 3, characterized in that The method further comprises: In response to the progress display instruction, display a channel list including each channel, and identify the channel of the video being urgently modeled in the channel list; In response to a selection instruction of the channel list, displaying a time period list, wherein the time period list includes a shooting time period of each video of the channel indicated by the selection instruction; If the channel indicated by the selection instruction is a channel of a video being urgently modeled, the method further includes: The shooting time period of the video being urgently modeled is identified in the time period list.

5. The method according to claim 2, characterized in that The method further comprises: In response to a continue search instruction during the process of modeling the first video, while at least continuing to model targets that have not yet been modeled in the first video, a first target is determined from the targets that have already been modeled from the first video, and at least a first image of each determined first target is displayed, wherein the first target is a target that matches the first text, and the first text includes a target category and features possessed by the target.

6. The method according to claim 1, characterized in that The first target is obtained by matching through a preset matching algorithm; The method further comprises: In response to a labeling instruction, determining a first image indicated by the labeling instruction as a labeling image, and determining a ground truth value labeled for the labeling image; Training the preset matching algorithm using the annotated image and the true value annotated for the annotated image; A new first target is obtained by re-matching using the trained preset matching algorithm, and the step of displaying the first image of at least one first target hit by the search condition is returned to be executed.

7. The method according to claim 1, characterized in that The displaying of the first image of at least one first target hit by the search condition includes: Displaying a first image of at least one first target hit by the search condition and displaying a first text; The method further comprises: In response to an edit instruction for the displayed first text, display a tab page for each feature of the category of the first target, wherein each feature tab page displays its own sub-features; In response to a fifth selection instruction of the displayed sub-feature, generating a second text including a third sub-text and a fourth sub-text as a new first text; wherein the third sub-text is used to describe all sub-features indicated by the fifth selection instruction, and the fourth sub-text is used to describe the category of the first target; A first target matching the new first text is determined, and the process returns to the step of displaying a first image of at least one first target hit by the search condition and displaying the first text.

8. The method according to claim 1, characterized in that The displaying of the first image of at least one first target hit by the search condition includes: Displaying the respective archive covers of some or all of the first targets hit by the search condition, wherein the archive represented by the archive cover of each first target includes at least one first image of the first target itself; in response to an expansion instruction, displaying some or all of the first images included in the archive represented by the archive cover indicated by the expansion instruction; and / or, All first images of at least one first object hit by the search condition are displayed respectively.

9. The method according to claim 1, characterized in that The method further comprises: Display multiple category cards representing different categories in the vicinity of the search box; In response to a target category selection instruction, determining a category card selected by the target category selection instruction, taking the category represented by the selected category card as a first category, and displaying respective tabs for each feature of the target in the first category, wherein each feature tab displays its own sub-features; In response to the second selection instruction of the displayed sub-feature, a first text including a first sub-text and a second sub-text is generated; wherein the first sub-text is used to describe all sub-features indicated by the second selection instruction, and the second sub-text is used to describe the first category.

10. The method according to claim 9, characterized in that The method further comprises: Displaying the first text in the search box, and displaying a mark and a modification control on the first text, wherein the mark is used to mark each feature sub-text, wherein the feature sub-text is a sub-text in the first text used to describe a sub-feature, and different feature sub-texts are used to describe different sub-features; In response to a third selection instruction, determining the feature subtext corresponding to the modification control indicated by the third selection instruction as the first feature subtext, and displaying a subfeature display interface, the subfeature display interface including other subfeatures on the same tab page as the subfeature described by the first feature subtext; In response to the displayed fourth selection instruction of the other sub-feature, the first feature subtext in the first text is replaced with a second feature subtext to obtain a new first text, wherein the second feature subtext is a subtext used to describe the other sub-feature indicated by the fourth selection instruction.

11. The method according to claim 1, wherein The first image is a sub-image of a first area where the first target is located in a first video frame, and the first video frame is a video frame in which the first target is detected in the first video; The method further comprises: In response to a zoom instruction for the displayed first image, a second image corresponding to the first image indicated by the zoom instruction is displayed, wherein the second image is a sub-image of a second area where the first target is located in the first video frame, and the first area is a proper subset of the second area.

12. The method according to claim 11, characterized in that The method further comprises: In response to a first playback instruction for a displayed first image, taking the first image indicated by the first playback instruction as a target image, determining a first video frame containing the target image as a target video frame, determining a sub-video of the first video that contains the target video frame and continuously detects the first target in the target image as a first target video, and displaying the first target video in a preset display window; and / or, In response to the second playback instruction of the displayed first image, the first image indicated by the second playback instruction is determined as the target image, the first video frame where the target image is located is determined as the target video frame, each first video frame in the first video that is continuous with the target video frame and detects the first target in the target image is obtained, each first image in each first video frame is obtained, and the second target video is obtained, and the second target video is played in a floating window at the position where the target image is displayed.

13. The method according to claim 11, characterized in that The displaying of the second image corresponding to the first image indicated by the zoom instruction includes: A second image corresponding to the first image indicated by the zoom-in instruction is displayed, and a target frame of the first target is displayed in the second image, where the first target is a target matching the first text.

14. The method according to claim 12, characterized in that The method further comprises: In response to the target search function enabling instruction, displaying a target frame of each detected target in the video frame currently displayed and played in the preset display window; In response to the target selection instruction, detecting in the first video targets that match the targets in the target frame indicated by the target selection instruction as fourth targets; The target information of each of the fourth targets is displayed respectively.

15. The method according to claim 1, wherein The displaying of a first image of at least one first target hit by the search condition includes: Filtering all first images of at least one static first target that are hit by the search condition to select all images that are hit by a preset filtering condition as filtered first images of the static first target; At least all filtered first images of the static first object are displayed.

16. The method according to claim 2, characterized in that The target in the first video is modeled in advance by: Extracting video frames from the first video according to a preset frame extraction interval, and modeling static objects appearing in the extracted video frames; and / or A video frame in which a dynamic object appears is determined from the first video, and a model is performed on the dynamic object appearing in the determined video frame.

17. A video processing device, characterized in that: The video processing device is used to implement the target detection method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Method and apparatus for video retrieval

    CN104798068A

  • Keyword searching method, device and equipment

    CN116627908A

  • Image retrieval method, device and system

    CN117743632A

  • Target retrieval method and device

    CN117743633A

  • Information processing apparatus, search method, and non-transitory computer readable medium storing program

    US20220179899A1