Intelligent image processing method, device and equipment for open self-defined detection target

By acquiring user-defined object description information and utilizing an open-vocabulary target detection model, the orientation of the image acquisition device is controlled to position the target of interest at a preset location. This solves the problem that existing image acquisition devices cannot customize the detection target, thus improving the usability and intelligence of the device.

CN120416668BActive Publication Date: 2025-11-04HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510897125.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-04
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing image acquisition equipment cannot customize detection targets according to users' individual needs, resulting in low usability and intelligence.

Method used

By acquiring user-defined object description information, a pre-trained open vocabulary object detection model is used to detect whether there is a target of interest in the image frame, and the orientation of the image acquisition device is controlled so that the target of interest is located in a preset position.

Benefits of technology

It enables continuous shooting of targets of interest in image frames according to user needs, improving the availability and intelligence of image acquisition equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416668B_ABST
    Figure CN120416668B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides an open intelligent image processing method and device for customizing a detection target, and relates to the technical field of computer vision. The method comprises the following steps: acquiring object description information input by a user in a customized manner; wherein the customized input manner comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; the object description information is used to determine a target of interest of the user; based on the object description information, it is detected whether the target of interest exists in a collected image frame; if the target of interest exists in the image frame, the direction of image collection is controlled so that the target of interest is located at a preset position in the image frame. In this way, the availability and intelligent degree of the image collection device are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and in particular to an intelligent image processing method, device and equipment for open self-defined detection targets. BACKGROUND

[0002] Generally, the field of view of a camera is fixed. When such a camera is used to implement image acquisition on a moving object, once the moving object is out of the fixed field of view of the camera, the camera cannot acquire images of the moving object. In order to achieve continuous and effective image acquisition on the moving object, the camera can be mounted on a holder, so that when the holder is controlled to rotate, the holder can drive the camera to rotate.

[0003] However, the type of target that the above image acquisition device supports for continuous shooting is usually preset at the factory, and cannot be continuously shot according to the individual needs of the user. For example, an image acquisition device is set to support continuous shooting of vehicles at the factory, and the image acquisition device only supports continuous shooting of vehicles, and cannot achieve continuous shooting of other targets (pedestrians and animals) other than vehicles. This results in low availability and intelligence of the image acquisition device. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide an intelligent image processing method, device and equipment for open self-defined detection targets, so as to improve the availability and intelligence of the image acquisition device. The specific technical solutions are as follows:

[0005] In a first aspect, the embodiments of the present application provide an intelligent image processing method for open self-defined detection targets, which comprises:

[0006] Obtaining object description information input by a user; wherein the input mode of the user includes at least one of the following: inputting text, inputting images, inputting audio, and triggering an interactive instruction containing position information in an image frame; and the object description information is used to determine a target of interest of the user;

[0007] Based on the object description information, detecting whether the target of interest exists in an acquired image frame;

[0008] If the target of interest is detected in the image frame, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frame.

[0009] In some embodiments, the data format of the object description information is one of text, image and audio;

[0010] The object description information is used to determine whether the image frame contains the target of interest.

[0011] If the image frame is the first image frame collected after the object description information is received, or if the target of interest is not detected in the continuous specified number of image frames before the image frame, an open-vocabulary target detection model is used to determine whether the image frame contains the target of interest based on the object description information and the image frame.

[0012] If the target of interest is detected in at least one of the continuous specified number of image frames, the open-vocabulary target detection model is used to detect the image frame, and the pixel position of the detected target and the pixel position of the target of interest detected in the continuous specified number of image frames are used to determine whether the image frame contains the target of interest.

[0013] In some embodiments, the open-vocabulary target detection model is used to detect the image frame, including:

[0014] The object description information and the image frame are input into a pre-trained open-vocabulary target detection model to obtain the pixel position of the target in the image frame that meets the object description information.

[0015] Alternatively,

[0016] The image frame and the preset prompt word are input into a pre-trained open-vocabulary target detection model to obtain the pixel position of the target in the image frame that meets the preset prompt word.

[0017] In some embodiments, the object description information is triggered by a user in a specified image frame.

[0018] The object description information is used to determine whether the image frame contains the target of interest.

[0019] The pixel position of each target in the specified image frame is obtained using a pre-trained open-vocabulary target detection model, and the pixel position of each target and the position information contained in the object description information are used to determine whether the image frame contains the target of interest.

[0020] For each image frame after the specified image frame, if it is detected that at least one of the continuous specified number of image frames before the image frame has the target of interest, pixel positions of each target in the image frame are obtained using the open vocabulary target detection model, and whether the target of interest exists in the image frame is determined according to the determined pixel positions of each target and the pixel positions of the target of interest detected in the continuous specified number of image frames.

[0021] In some embodiments, if it is detected that the target of interest exists in the image frame, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frame, including:

[0022] If the target of interest exists in the image frame obtained at the first detection time, a moving speed of the target of interest in the direction of the imaging plane at the first detection time is calculated as a first moving speed based on the pixel position of the target of interest in the image frame.

[0023] A moving speed of the target of interest in the direction of the imaging plane at a third detection time is calculated as a third moving speed according to a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time and the first moving speed, wherein the third detection time is a next detection time of the second detection time, and the second detection time is a next detection time of the first detection time.

[0024] When the second detection time is reached, the speed of the motor is adjusted according to the third moving speed, the current speed of the motor in the image acquisition device at the current time, and the maximum speed supported by the motor, so that the rotating speed of the lens in the image acquisition device at the third detection time is consistent with the third moving speed, and the moving distance of the target of interest in the direction of the imaging plane from the second detection time to the third detection time is consistent with the rotating angle of the lens.

[0025] In some embodiments, calculating the moving speed of the target of interest in the direction of the imaging plane at the third detection time as the third moving speed according to the second moving speed of the target of interest in the direction of the imaging plane at the historical detection time before the first detection time and the first moving speed includes:

[0026] Determining a pixel distance between the target of interest and a reference stationary object in the image frame.

[0027] input the determined pixel distance, a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time, and the first moving speed into a pre-trained speed prediction model to obtain a moving speed of the target of interest in the direction of the imaging plane at the third detection time as a third moving speed;

[0028] The speed prediction model is trained based on sample data and corresponding sample labels.

[0029] The sample data includes a pixel distance between a sample object and a reference stationary object in a sample image frame collected at a detection time, a moving speed of the sample object in the direction of the imaging plane at a historical detection time before the detection time, and a moving speed of the sample object in the direction of the imaging plane at the detection time. The sample label corresponding to one sample data includes a speed of the sample object at a second detection time after the detection time.

[0030] In some embodiments, the first moving speed is calculated based on the pixel position of the target of interest in the image frame.

[0031] In a case where an image quality of an image region occupied by the target of interest in the image frame meets a preset condition, the first moving speed is calculated based on the pixel position of the target of interest in the image frame.

[0032] The preset condition includes at least one of the following: the image region is not overexposed, the image region is not too dark, and the image region is clear.

[0033] In some embodiments, when the second detection time is reached, the speed of the motor is adjusted according to the third moving speed, a current speed of the motor in the image acquisition device at the current time, and a maximum speed supported by the motor, so that a rotation speed of a lens in the image acquisition device at the third detection time is consistent with the third moving speed, and a moving distance of the target of interest in the direction of the imaging plane from the second detection time to the third detection time is consistent with a rotation angle of the lens.

[0034] When the second detection time is reached, a peak speed required by the motor to be reached at a specified time between the second detection time and the third detection time is determined according to the third moving speed, a current speed of the motor in the image acquisition device at the current time, and a maximum speed supported by the motor.

[0035] The speed of the motor at the current time is taken as an initial speed, and the speed of the motor is adjusted in an increasing or decreasing manner so that the speed of the motor reaches the peak speed at the specified time.

[0036] When the specified time is reached, if the speed of the motor at the current time does not reach the supported maximum speed, the speed of the motor at the current time is taken as an initial speed, and the speed of the motor is adjusted in an increasing or decreasing manner so that the speed of the motor reaches the third moving speed at the third detection time.

[0037] The method further comprises:

[0038] When the specified time is reached, if the speed of the motor at the current time is the supported maximum speed, the motor is controlled to maintain the speed at the current time until the third detection time.

[0039] In some embodiments, when the second detection time is reached, the peak speed required by the motor at the specified time between the second detection time and the third detection time is determined according to the third moving speed, the speed of the motor at the current time in the image acquisition device, and the maximum speed supported by the motor, comprising:

[0040] When the second detection time is reached, the peak speed required by the motor at the specified time between the second detection time and the third detection time is determined according to the third moving speed, the speed of the motor at the current time in the image acquisition device, and the maximum speed supported by the motor according to a preset formula.

[0041] The preset formula is represented as:

[0042]

[0043] The peak speed is represented as ωp; The maximum speed supported by the motor is represented as ωmax; The third moving speed is represented as v3; The speed of the motor at the second detection time is represented as ω2.

[0044] The second aspect of the embodiments of the present application provides an open self-defined detection target intelligent image processing device, which comprises:

[0045] A lens is used to acquire image frames.

[0046] A processor is used to execute the open self-defined detection target intelligent image processing method.

[0047] In some embodiments, the device further comprises a gimbal;

[0048] The processor is specifically configured to send a control instruction to the gimbal.

[0049] The gimbal is configured to adjust an image acquisition direction of the lens according to the control instruction sent by the processor, so that the target of interest is located at a preset position in an image frame.

[0050] In a third aspect, the embodiments of the present application provide an intelligent image processing system for open self-defined detection target, which comprises:

[0051] A client is configured to receive object description information input by a user in a self-defined manner; the self-defined manner comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; the object description information is used to determine a target of interest of the user.

[0052] An image processing end is configured to acquire an image frame, and detect whether the target of interest exists in the acquired image frame based on the object description information; if the target of interest exists in the image frame, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frame.

[0053] In some embodiments, the client is integrated in the image processing end.

[0054] In some embodiments, the system further comprises a cloud platform.

[0055] The cloud platform is configured to receive the object description information sent by the client and forward the object description information to the image processing end.

[0056] In a fourth aspect, the embodiments of the present application provide an intelligent image processing system for open self-defined detection target, which comprises:

[0057] A client is configured to receive object description information input by a user in a self-defined manner; the self-defined manner comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; the object description information is used to determine a target of interest of the user.

[0058] An image processing end is configured to acquire an image frame.

[0059] A cloud platform is configured to receive the object description information sent by the client and receive an image frame sent by the image processing end; and detect whether the target of interest exists in the image frame based on the object description information, and send a detection result to the image processing end.

[0060] The image processing end is further configured to control a direction of image acquisition according to the detection result sent by the cloud platform, so that the target of interest is located at a preset position in an image frame.

[0061] In a fifth aspect, the present application provides an intelligent image processing device for open self-defined detection target, comprising:

[0062] An object description information acquisition module is configured to acquire object description information input by a user in a self-defined manner, wherein the self-defined input manner comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; and the object description information is used to determine a target of interest of the user.

[0063] A detection module is configured to detect whether the target of interest exists in an acquired image frame based on the object description information.

[0064] A first control module is configured to control a direction of image acquisition if the target of interest is detected in the image frame, so that the target of interest is located at a preset position in an image frame.

[0065] In some embodiments, a data format of the object description information is one of text, an image, and audio.

[0066] The detection module comprises:

[0067] A first detection sub-module is configured to, for an acquired image frame, if it is detected that the target of interest does not exist in a continuous specified number of image frames before the image frame, or the image frame is the first image frame acquired after the object description information is received, detect whether the target of interest exists in the image frame based on the object description information and the image frame by using a pre-trained open-vocabulary target detection model.

[0068] A second detection sub-module is configured to, if it is detected that the target of interest exists in at least one of the continuous specified number of image frames, detect the image frame by using the open-vocabulary target detection model, and determine whether the target of interest exists in the image frame according to a pixel position of the detected target and a pixel position of the target of interest detected in the continuous specified number of image frames.

[0069] In some embodiments, the second detection sub-module is specifically configured to:

[0070] input the object description information and the image frame into a pre-trained open-vocabulary target detection model to obtain a pixel position of a target in the image frame that meets the object description information.

[0071] or,

[0072] inputting the image frame and the preset prompt word into a pre-trained open-vocabulary object detection model to obtain a pixel position of an object in the image frame that matches the preset prompt word.

[0073] In some embodiments, the object description information is triggered by a user in a specified image frame;

[0074] The detection module comprises:

[0075] The third detection submodule is configured to obtain a pixel position of each object in the specified image frame by using the pre-trained open-vocabulary object detection model, and determine whether the target of interest exists in the image frame according to the determined pixel position of each object and the position information contained in the object description information.

[0076] The fourth detection submodule is configured to, for each image frame after the specified image frame, if it is detected that the target of interest exists in at least one image frame among the continuous specified number of image frames before the image frame, obtain a pixel position of each object in the image frame by using the open-vocabulary object detection model, and determine whether the target of interest exists in the image frame according to the determined pixel position of each object and the pixel position of the target of interest detected in the continuous specified number of image frames.

[0077] In some embodiments, the first control module comprises:

[0078] The first speed calculation submodule is configured to, if the target of interest exists in the image frame obtained at the first detection time, calculate a moving speed of the target of interest in the direction of the imaging plane at the first detection time as a first moving speed based on the pixel position of the target of interest in the image frame.

[0079] The second speed calculation submodule is configured to calculate a moving speed of the target of interest in the direction of the imaging plane at a third detection time as a third moving speed according to a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time and the first moving speed; the third detection time is a next detection time of the second detection time; and the second detection time is a next detection time of the first detection time.

[0080] a control submodule configured to, when the second detection time is reached, adjust a rotation speed of the motor according to the third moving speed, a current rotation speed of the motor at the third detection time, and a maximum rotation speed supported by the motor, so that a rotation speed of a lens in the image acquisition device at the third detection time is consistent with the third moving speed, and a moving distance of the target of interest in a direction of an imaging plane from the second detection time to the third detection time is consistent with a rotation angle of the lens.

[0081] In some embodiments, the second speed calculation submodule is specifically configured to:

[0082] determine a pixel distance between the target of interest and a reference stationary object in the image frame;

[0083] input the determined pixel distance, a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time, and the first moving speed into a pre-trained speed prediction model to obtain a moving speed of the target of interest in the direction of the imaging plane at the third detection time as the third moving speed;

[0084] wherein the speed prediction model is trained based on sample data and corresponding sample labels;

[0085] the sample data includes a pixel distance between a sample object and a reference stationary object in a sample image frame collected at a detection time, a moving speed of the sample object in the direction of the imaging plane at a historical detection time before the detection time, and a moving speed of the sample object in the direction of the imaging plane at the detection time; and a sample label corresponding to one sample data includes a speed of the sample object at a second detection time after the detection time.

[0086] In some embodiments, the first speed calculation submodule is specifically configured to:

[0087] in a case where an image quality of an image region occupied by the target of interest in the image frame meets a preset condition, calculate a moving speed of the target of interest in the direction of the imaging plane at the first detection time based on a pixel position of the target of interest in the image frame as the first moving speed;

[0088] wherein the preset condition includes at least one of the following: the image region is not overexposed, the image region is not too dark, and the image region is clear.

[0089] In some embodiments, the control submodule is specifically configured to:

[0090] a peak speed determination unit configured to determine a peak speed required by the motor at a specified time between the second detection time and the third detection time according to the third moving speed, a current speed of the motor at the current time, and a maximum speed supported by the motor when the second detection time is reached;

[0091] a first speed adjustment unit configured to adjust the speed of the motor in an increasing or decreasing manner with the current speed of the motor at the current time as an initial speed, so that the speed of the motor reaches the peak speed at the specified time;

[0092] a second speed adjustment unit configured to adjust the speed of the motor in an increasing or decreasing manner with the current speed of the motor at the current time as an initial speed, so that the speed of the motor reaches the third moving speed at the third detection time when the specified time is reached and the current speed of the motor at the current time is not the maximum speed supported by the motor;

[0093] The apparatus further includes:

[0094] a second control module configured to control the motor to keep the current speed until the third detection time when the specified time is reached and the current speed of the motor at the current time is the maximum speed supported by the motor.

[0095] In some embodiments, the peak speed determination unit is specifically configured to:

[0096] determine the peak speed according to the third moving speed, the current speed of the motor at the current time, and the maximum speed supported by the motor according to a preset formula when the second detection time is reached;

[0097] wherein the preset formula is represented as:

[0098]

[0099] denotes the peak speed; denotes the maximum speed supported by the motor; denotes the third moving speed; denotes the speed of the motor at the second detection time.

[0100] In yet another aspect of the embodiments of the present application, an electronic device is provided, which includes:

[0101] a memory configured to store a computer program;

[0102] A processor is configured to implement the open self-defined detection target intelligent image processing method.

[0103] In another aspect of the embodiments of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the open self-defined detection target intelligent image processing method.

[0104] In another aspect of the embodiments of the present application, a computer program product is provided, and the computer program product contains instructions. When the computer program product is run on a computer, the computer is caused to execute the open self-defined detection target intelligent image processing method.

[0105] The embodiments of the present application have the following beneficial effects:

[0106] Based on the open self-defined detection target intelligent image processing method provided by the embodiments of the present application, the object description information for determining the target of interest of the user can be input in a self-defined manner. Correspondingly, whether the target of interest exists in the collected image frame can be detected based on the object description information. Further, if the target of interest exists in the image frame, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frame. That is, the electronic device can control the image acquisition device to continuously shoot the target (i.e., the target of interest) in the image frame that meets the object description information according to the actual needs of the user, thereby improving the usability and intelligent degree of the image acquisition device.

[0107] Of course, implementing any product or method of the present application does not necessarily need to achieve all the advantages described above at the same time. BRIEF DESCRIPTION OF DRAWINGS

[0108] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other embodiments can also be obtained by those skilled in the art based on these drawings.

[0109] Figure 1 The first flowchart of the open self-defined detection target intelligent image processing method provided by the embodiments of the present application;

[0110] Figure 2 The first flowchart of the open self-defined detection target intelligent image processing method provided by the embodiments of the present application;

[0111] Figure 3 The second flowchart of the open self-defined detection target intelligent image processing method provided by the embodiments of the present application;

[0112] Figure 4 A third flowchart for controlling after detecting a target of interest is provided for the embodiments of the present application.

[0113] Figure 5 A flowchart schematic diagram of image acquisition device control is provided for the embodiments of the present application.

[0114] Figure 6 A structural schematic diagram of an intelligent image processing device is provided for the embodiments of the present application.

[0115] Figure 7 A structural schematic diagram of an intelligent image processing system is provided for the embodiments of the present application.

[0116] Figure 8 A structural schematic diagram of another intelligent image processing system is provided for the embodiments of the present application.

[0117] Figure 9 A structural diagram of an open self-defined target intelligent image processing device is provided for the embodiments of the present application.

[0118] Figure 10 A structural schematic diagram of an electronic device is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0119] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art based on the present application belong to the scope of protection of the present application.

[0120] In order to solve the technical problem of low availability and intelligence of image acquisition devices, the present application provides an open self-defined target intelligent image processing method, which can be applied to an electronic device.

[0121] Among them, the description of the open self-defined target intelligent image processing device can be seen in the subsequent embodiments.

[0122] Reference Figure 1 , Figure 1 A first flowchart of the open self-defined target intelligent image processing method provided for the embodiments of the present application, the method comprises the following steps:

[0123] S101: Obtain the object description information input by the user.

[0124] The self-defined input manner includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; and the object description information is used to determine the target of interest of the user.

[0125] S102: Based on the object description information, it is detected whether the target of interest exists in the collected image frame.

[0126] S103: If it is detected that the target of interest exists in the image frame, the direction of image collection is controlled, so that the target of interest is located at a preset position in the image frame.

[0127] Based on the open self-defined detection target intelligent image processing method provided in the embodiments of the present application, the user can input the object description information used to determine the target of interest of the user in a self-defined input manner. Correspondingly, it can be detected whether the target of interest exists in the collected image frame based on the object description information. Further, if it is detected that the target of interest exists in the image frame, the direction of image collection is controlled, so that the target of interest is located at a preset position in the image frame. That is, the electronic device can control the image collection device to continuously shoot the target (i.e., the target of interest) meeting the object description information in the image frame according to the actual needs of the user, and improve the usability of the image collection device.

[0128] For step S101, the user can input the object description information used to determine the target of interest in a self-defined input manner.

[0129] The self-defined input manner includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame.

[0130] In an implementation manner, the electronic device can provide an interactive interface used to receive the object description information input by the user. Alternatively, a client independent of the electronic device can provide an interactive interface used to receive the object description information input by the user, and the client can communicate with the electronic device. Correspondingly, the user can input the object description information through the interactive interface of the client, and then the client can send the object description information input by the user to the electronic device.

[0131] The self-defined input manner includes:

[0132] Manner one, inputting text.

[0133] In this manner, the user can input the text representing the target to be continuously shot in the text box in the interactive interface, click a submit button, and then the electronic device can acquire the text input by the user as the object description information. That is, the data format of the object description information is text.

[0134] For example, the target to be continuously photographed is a person wearing a white coat, and the user can input the text "person wearing a white coat".

[0135] Method two, input an image.

[0136] The user can upload an image containing the target to be continuously photographed through an image uploading component in the interactive interface, and then the electronic device can obtain the image input by the user as the object description information. That is, the data format of the object description information is an image.

[0137] For example, the target to be continuously photographed is a lynx cat, and the user can input an image containing the lynx cat.

[0138] Method three, input audio.

[0139] The user can upload audio representing the target to be continuously photographed through an audio uploading component in the interactive interface, and then the electronic device can obtain the audio input by the user as the object description information. At this time, the data format of the object description information is audio.

[0140] For example, the target to be continuously photographed is a vehicle of a certain brand, and the user can upload audio, and the content of the audio is "vehicle of a certain brand".

[0141] Method four, trigger an interactive instruction containing position information in an image frame.

[0142] The electronic device can display an image frame captured by the image capturing device in real time in the interactive interface. If the target to be continuously photographed by the user exists in the currently displayed image frame, the user can trigger an interactive instruction containing position information in the image frame by clicking or sliding selection at the position of the target to be continuously photographed in the image frame. Correspondingly, the electronic device can obtain the interactive instruction containing position information triggered by the user in the image frame.

[0143] It can be understood that since the target to be continuously photographed can be in a moving state, and correspondingly, the position of the same target in image frames captured at different times can also be different, the object description information input by the user through the triggering of the interactive instruction is only for the image frame to which the interactive instruction is directed (which can be referred to as the specified image frame of the interactive instruction). That is, the electronic device can only detect whether the target of interest exists in the specified image frame of the interactive instruction based on the interactive instruction. It cannot detect whether the target of interest exists in other image frames except the specified image frame of the interactive instruction based on the interactive instruction.

[0144] For step S102, the electronic device can detect whether the target of interest exists in the image frame captured at the moment when the object description information input by the user is obtained.

[0145] For any collected image frame, if the detection result of the image frame indicates that the image frame contains the object of interest, the electronic device can further determine the pixel position of the object of interest in the image frame.

[0146] When the object (i.e., the object of interest) conforming to the object description information enters the field of view of the image collection device, the electronic device can start continuous shooting of the object of interest. That is, the electronic device can control the direction of image collection so that the object of interest is located at a preset position in the image frame. The specific process of controlling the direction of image collection will be described in subsequent embodiments.

[0147] It can be understood that, since the range that can be adjusted by the direction of image collection is limited, the maximum field of view that can be collected by the image collection device is also limited. For an object of interest in a motion state, as the object of interest continuously moves, the object of interest can leave the maximum field of view that can be collected by the image collection device. Accordingly, the image collection device can stop continuous shooting of the object of interest.

[0148] In an implementation manner, after the image collection device stops continuous shooting of the object of interest, the electronic device can adjust the direction of image collection of the image collection device to the initial direction. The electronic device can adjust the direction of image collection of the image collection device by controlling the rotation of the motor in the image collection device. The specific process of adjusting the direction of image collection of the image collection device will be described in subsequent embodiments.

[0149] If the data format of the object description information is one of text, image, and audio, after an object of interest leaves the maximum field of view that can be collected by the image collection device for a period of time, the image collection device can re-detect whether there is an object (a new object of interest) conforming to the object description information in the collected image frame, and perform continuous shooting of the new object of interest after detecting the new object of interest.

[0150] In an implementation manner, after each new object of interest is determined, the electronic device can assign a unique identifier (ID) corresponding to the object of interest to the object of interest. That is, the identifiers corresponding to different objects of interest are different. For example, the unique identifier corresponding to an object of interest can represent the time sequence in which the object of interest appears in the image collection device. Subsequently, the user can refer to the video clips corresponding to each object of interest through the unique identifier corresponding to each object of interest.

[0151] In some embodiments, the data format of the object description information is one of text, image, and audio.

[0152] Correspondingly, step S102 includes:

[0153] Step S1021: For one image frame collected, if it is detected that there is no target of interest in the continuous specified number of image frames before the image frame, or the image frame is the first image frame collected after receiving the object description information, then based on the object description information and the image frame, the pre-trained open vocabulary target detection model is used to detect whether there is a target of interest in the image frame.

[0154] Step S1022: If it is detected that there is at least one image frame with a target of interest in the continuous specified number of image frames, then the open vocabulary target detection model is used to detect the image frame; and whether there is a target of interest in the image frame is determined according to the pixel position of the detected target and the pixel position of the target of interest detected in the continuous specified number of image frames.

[0155] In the embodiments of the present application, for step S1021, for one image frame collected, if the image frame is the first image frame collected after receiving the object description information, it indicates that the electronic device needs to detect whether there is a target of interest in the field of view of the image collection device from the image frame. Correspondingly, based on the obtained object description information and the first image frame, the pre-trained open vocabulary target detection model can be used to detect whether there is a target of interest in the image frame.

[0156] Alternatively, for one image frame collected, if it is detected that there is no target of interest in the continuous specified number of image frames before the image frame, it indicates that in the historical time period corresponding to the continuous specified number of image frames before the image frame, no target of interest has been detected to appear in the field of view of the image collection device. This situation may be since the beginning (i.e., since the moment of receiving the object description information) no detection has been made, or it may be that the image frame of the historical detection has a target of interest, and as the target of interest moves, the target of interest leaves the field of view of the image collection device. Since in the historical time period corresponding to the continuous specified number of image frames before the image frame, no target of interest has been detected to appear in the field of view of the lens, at this time, the electronic device can re-determine the target of interest. That is, the electronic device can only use the pre-trained open vocabulary target detection model based on the obtained object description information and the image frame to detect whether there is a target of interest in the image frame.

[0157] Wherein, the continuous specified number can be determined by the technical personnel according to the actual demand, for example, the continuous specified number can be 10.

[0158] Correspondingly, the electronic device can take the obtained object description information as a prompt word (Prompt) required to input to the open vocabulary object detection model. That is, the image frame and the prompt word (i.e., the object description information) are input to the pre-trained open vocabulary object detection model to obtain a detection result indicating whether the target of interest exists in the image frame.

[0159] The pre-trained open vocabulary object detection model (OVOD, Open Vocabulary Object Detection) is a target detection model combined with a large language model (LLM, Large Language Model). For example, the open vocabulary object detection model can be YOLO-World (You Only Look Once-World, real-time open vocabulary object detection model), or DINO (Detection Transformer, target detection model based on a transformer network), etc.

[0160] For step S1022, for the collected image frame, if it is detected that at least one image frame among the continuous specified number of image frames before the image frame has the target of interest, it indicates that the historical time period corresponding to the specified number of image frames before the image frame is collected has detected that the target of interest appears in the field of view range of the lens.

[0161] Correspondingly, in order to realize continuous shooting of the same target of interest as much as possible, the electronic device can use the open vocabulary object detection model to detect the image frame to obtain the pixel position of the detected target (which can be referred to as an alternative target). Further, in combination with the pixel position of the target of interest detected from the continuous specified number of image frames before the image frame, the alternative target in the image frame is determined as the target of interest.

[0162] In some embodiments, the pixel position of the alternative target in the image frame can be obtained according to any one of the following mode a or mode b:

[0163] Mode a: input the object description information and the image frame to the pre-trained open vocabulary object detection model to obtain the pixel position of the target in the image frame that meets the object description information.

[0164] In the embodiments of the present application, the electronic device can input the object description information and the image frame to the pre-trained open vocabulary object detection model to obtain the pixel position of all targets (which can be referred to as first alternative targets) in the image frame that meet the object description information. The specific detection process using the pre-trained open vocabulary object detection model can refer to the above embodiments, which will not be repeated here.

[0165] Based on the above processing, it can be ensured that the detected target (the first alternative target) meets the object description information input by the user. Subsequently, determining the target of interest from the first alternative target can improve the accuracy of the determined target of interest.

[0166] Option b: input the image frame and the preset prompt word into a pre-trained open vocabulary target detection model to obtain the pixel position of the target in the image frame that meets the preset prompt word.

[0167] In the embodiments of the present application, the electronic device can input the image frame and the preset prompt word into a pre-trained open vocabulary target detection model to obtain the pixel position of the target (which can be referred to as the second alternative target) in the image frame that meets the preset prompt word.

[0168] The preset prompt word can be a prompt word set by a technician in advance. For example, the preset prompt word can include at least one of the following: a person, an animal, a vehicle, and the like.

[0169] Based on the above processing, the preset prompt word can be a general prompt word set by a technician in advance, which can ensure that various movable targets in the image frame can be detected as much as possible, avoiding missing detection of targets.

[0170] It can be understood that, in the plurality of continuous image frames (the image frame and the continuous specified number of image frames before the image frame), for the same target, the positions of the target in adjacent two image frames are usually close, and the position of the target in the plurality of continuous image frames is usually continuous.

[0171] Therefore, after obtaining the pixel position of the alternative target detected in the image frame, the electronic device can determine whether the target of interest exists in the image frame based on the pixel position of each alternative target and the pixel position of the target of interest detected in the continuous specified number of image frames (which can be referred to as the historical pixel position of the target of interest).

[0172] In an implementation manner, the electronic device can use a preset object tracking (OT) algorithm to determine the pixel position of the target of interest in the image frame based on the pixel position of each alternative target and the pixel position of the target of interest in the collected historical image frames (i.e., the pixel position of the target of interest detected in the continuous specified number of image frames). For example, the object tracking algorithm can be a Hungarian matching algorithm.

[0173] Correspondingly, the electronic device can obtain the pixel positions of the target of interest in the historical image frames. For the pixel position of each candidate target in the image frame, the intersection over union (IOU) between the image region occupied by the candidate target in the image frame and the image region occupied by the target of interest in the historical image frames is calculated to obtain the cost matrix corresponding to the candidate target in the image frame. Wherein, the greater the IOU corresponding to the candidate target in the image frame, the smaller the value in the corresponding cost matrix. Further, among the candidate targets in the image frame, the candidate target corresponding to the minimum total cost represented by the cost matrix and not greater than a preset threshold is determined as the target of interest in the image frame, and the pixel position of the target of interest in the image frame is determined. If there is no candidate target corresponding to the cost matrix not greater than the preset threshold, it can be determined that none of the candidate targets in the image frame is the target of interest in the historical image frames, and further, it can be determined that there is no target of interest in the image frame.

[0174] In another implementation, the electronic device can use a pre-trained target following model to determine the pixel position of the target of interest in the image frame according to the pixel positions of the candidate targets in the image frame and the pixel positions of the target of interest in the historical image frames.

[0175] For example, the target following model can be a deep learning model trained based on a DeepSORT (Deep Simple Online and Realtime Tracking) algorithm.

[0176] Correspondingly, the electronic device can input the image frame, the pixel positions of the candidate targets in the image frame, the historical image frames, and the pixel positions of the target of interest in the historical image frames into the pre-trained target following model.

[0177] Correspondingly, the multi-label following model can extract the image features of the image regions occupied by the candidate targets in the image frame and the image features of the image regions occupied by the target of interest in the historical image frames. For each candidate target in the image frame, the IOU between the image region occupied by the candidate target and the image region occupied by the target of interest in the historical image frames is calculated. Further, the similarity between the image features of the image regions occupied by the candidate targets and the image features of the image regions occupied by the target of interest in the historical image frames is combined with the calculated IOUs to determine whether the target of interest exists in the candidate targets in the image frame.

[0178] Thus, the target following model performs multimodal matching between the image features and pixel positions of each candidate target in the image frame and the image features and pixel positions of the target of interest in each historical image frame, so that the robustness of the matching can be improved, and the method is especially suitable for a case where the candidate targets in the image frame are occluded. In a case where it is detected that the target of interest exists in at least one of the continuous specified number of image frames, the electronic device can determine whether the target of interest exists in the image frame to be processed according to the pixel positions of each candidate target in the image frame and the pixel positions of the target of interest in the historical image frames collected. Thus, the positioning of the target of interest in the image frame to be processed is realized, and the continuity of the detection of the same target is ensured.

[0179] Based on the above processing, when the data format of the object description information is one of text, image, and audio, the electronic device can determine the target of interest appearing in different time periods. That is, the electronic device can detect whether a new target of interest exists in the image frames collected after a target of interest leaves the field of view of the image acquisition device for a period of time. In addition, the user can customize the description information of the target of interest according to actual needs, and use the open vocabulary target detection model to detect an object in the image frame that matches the description information as the target of interest. Thus, the category of the object that can be continuously photographed can be expanded, and the individual needs of the user in different scenarios can be met.

[0180] If the object description information is triggered by the user in the specified image frame, it indicates that the object description information is only used to describe the target of interest in the specified image frame.

[0181] In some embodiments, the object description information is triggered by the user in the specified image frame. Accordingly, step S102 includes:

[0182] Step S102a: obtaining the pixel positions of each target in the specified image frame by using the pre-trained open vocabulary target detection model, and determining whether the target of interest exists in the image frame according to the determined pixel positions of each target and the position information contained in the object description information.

[0183] Step S102b: for each image frame after the specified image frame, if it is detected that the target of interest exists in at least one of the continuous specified number of image frames before the image frame, obtaining the pixel positions of each target in the image frame by using the open vocabulary target detection model, and determining whether the target of interest exists in the image frame according to the determined pixel positions of each target and the pixel positions of the target of interest detected in the continuous specified number of image frames.

[0184] In the embodiments of the present application, the user triggers an interaction instruction containing position information in an image frame (i.e., a specified image frame), and accordingly, the electronic device can determine a target of interest in the specified image frame by using the position information contained in the interaction instruction.

[0185] For step S102a, the electronic device can input the image frame and the preset prompt word into the pre-trained open vocabulary target detection model to obtain the pixel position of the target in the image frame that meets the preset prompt word.

[0186] The preset prompt word can be a prompt word pre-set by a technician. For example, the preset prompt word can include at least one of the following: a person, an animal, a vehicle, and the like.

[0187] In this way, the preset prompt word can be a general-purpose prompt word pre-set by a technician, which can ensure that various movable targets in the image frame can be detected as much as possible, avoiding missing detection of targets.

[0188] It can be understood that if the position information contained in the interaction instruction represents a pixel coordinate, the electronic device can determine, as the target of interest, the target in the detected targets whose image region contains the pixel coordinate.

[0189] If the pixel coordinate does not belong to the image region of any detected target, the electronic device can determine, as the target of interest, the target in the specified image frame that is closest to the pixel coordinate. The distance between a target and the pixel coordinate can be represented as the closest distance between each pixel coordinate in the image region occupied by the target and the pixel coordinate.

[0190] If the position information contained in the interaction instruction represents a block of image region in the specified image frame, the electronic device can determine, as the target of interest, the target in the specified image frame that has the highest intersection-over-union ratio with the image region. The way of calculating the intersection-over-union ratio can refer to the above embodiments, which will not be described here.

[0191] For step S102b, for each image frame after the specified image frame, the process of determining whether the target of interest exists in the image frame can refer to the process of detecting whether at least one of the continuous specified number of image frames before the image frame contains the target of interest, which will not be described here.

[0192] Based on the above processing, the user can determine the position information of the target of interest in the image to be processed by clicking or selecting a box. Furthermore, the electronic device can determine the target of interest required for continuous shooting from the image frame by using the position information, so that the operation steps of the user can be simplified and the user experience can be improved.

[0193] For step S103, for a collected image frame, if it is detected that the image frame contains the target of interest, the electronic device can control the direction of image acquisition, so that the target of interest is located at a preset position in the image frame.

[0194] For example, the preset position can be the position of the center point in the image frame.

[0195] It can be understood that the frequency of image acquisition by the image acquisition device is usually high, and the difference between adjacent two collected image frames is small. In order to reduce the calculation amount of the target, the electronic device can perform the detection process of steps S101-S102 on each image frame collected at a preset detection frequency.

[0196] In this way, detection is avoided for each collected image frame, and the demand pressure on the calculation resources of the electronic device is reduced.

[0197] Correspondingly, for a detection time, an image frame collected at the detection time (which can be referred to as a to-be-processed image frame). If the to-be-processed image frame contains the target of interest, the detection time corresponding to the to-be-processed image frame (i.e. the first detection time) is the first detection time in the process of continuous shooting of the target of interest by the image acquisition device.

[0198] Referring to Figure 2 , Figure 2 The first flowchart provided by the embodiment of the present application for controlling after detecting the target of interest.

[0199] S201: If the image frame obtained at the first detection time contains the target of interest, the moving speed of the target of interest in the direction of the imaging plane at the first detection time is calculated based on the pixel position of the target of interest in the image frame, as the first moving speed.

[0200] S202: According to the second moving speed of the target of interest in the direction of the imaging plane at the historical detection time before the first detection time and the first moving speed, the moving speed of the target of interest in the direction of the imaging plane at the third detection time is calculated, as the third moving speed.

[0201] Wherein, the third detection time is the next detection time of the second detection time; the second detection time is the next detection time of the first detection time.

[0202] S203: when the second detection time is reached, adjusting the rotation speed of the motor according to the third moving speed, the current rotation speed of the motor at the time, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection time is consistent with the third moving speed, and the moving distance of the target of interest in the direction of the imaging plane from the second detection time to the third detection time is consistent with the rotation angle of the lens.

[0203] Based on the above processing, the first detection time, the second detection time and the third detection time are three consecutive detection times. If there is a target of interest in the image frame obtained at the first detection time (which can be referred to as a to-be-processed image frame), the first moving speed of the target of interest in the direction of the imaging plane at the first detection time can be calculated based on the pixel position of the target of interest in the to-be-processed image frame. Similarly, for the historical detection time before the first detection time, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection time can also be calculated. Furthermore, the third moving speed of the target of interest in the direction of the imaging plane at the third detection time can be calculated according to the trend of the moving speed of the target of interest at different detection times. Accordingly, when the second detection time is reached, the rotation speed of the motor can be adjusted according to the third moving speed, the current rotation speed of the motor at the time, and the maximum rotation speed supported by the motor. So that the rotation speed of the lens in the image acquisition device at the third detection time is consistent with the third moving speed, that is, the lens rotation speed is adjusted from the second detection time, so that at the third detection time, the pixel position change speed of the target of interest in the image frame caused by the rotation of the lens in the image acquisition device is consistent with the pixel position change speed of the target of interest in the image frame caused by the movement of the target of interest itself, which can ensure the synchronization of the rotation speed of the lens in the image acquisition device and the moving speed of the target of interest, that is, the lens rotation speed can be adjusted according to the moving speed of the target of interest, so that the lens in the image acquisition device can continuously capture the moving target of interest. And from the second detection time to the third detection time, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens, that is, within the time interval from the second detection time to the third detection time, the pixel position of the target of interest in the image caused by the rotation of the lens can be used to compensate for the pixel position of the target of interest in the image caused by the movement of the target of interest itself, to ensure the synchronization of the rotation amount of the lens in the image acquisition device and the moving amount of the target of interest, that is, the moving amount of the target of interest is compensated by the rotation amount of the lens, so that the position of the moving target of interest in the image frame captured within the time interval is stable. In this way, the image acquisition device can continuously capture the target of interest to a certain extent.

[0204] For step S201, the first detection time point represents a detection time point at which the image frame obtained at the detection time point contains the target of interest.

[0205] After starting to continuously capture the target of interest by using the image acquisition device, the electronic device performs steps S201-S202 once at each detection time point to obtain the moving speed of the target of interest in the direction of the imaging plane at a second detection time point after the detection time point. The target of interest is an object that needs to be continuously captured by the image acquisition device, and the specific execution process of steps S201-S202 will be described in subsequent embodiments. For ease of description, the process of performing steps S201-S202 once by the electronic device can be referred to as a speed prediction process.

[0206] The time interval (which can be referred to as a detection time interval) between every two adjacent detection time points is preset by a technician according to the execution time required for performing the speed prediction process once. For example, the detection time interval is usually not less than the time consumed by the electronic device for performing the speed prediction process once according to steps S201-S202. For example, if the execution time required for performing the speed prediction process once is about 120 milliseconds (i.e., 0.12 seconds), the detection time interval can be set to 120 milliseconds.

[0207] It can be understood that the detection time point and the time point at which the image acquisition device captures the image frame can not be synchronized. That is, the frame interval between two adjacent image frames captured by the image acquisition device can be different from the length of the time interval between every two adjacent detection time points.

[0208] The frame interval between two adjacent image frames captured by the image acquisition device is determined by the frame rate of the image acquisition device. For example, when the frame rate of the image acquisition device is 50 fps (frames per second), the frame interval between two adjacent image frames captured by the image acquisition device is 0.02 seconds.

[0209] In an implementation manner, when a detection time point is reached, if the image acquisition device captures an image frame at the detection time point, the electronic device can determine that the image frame is the image frame captured at the detection time point.

[0210] Alternatively, when a detection time point is reached, if the image acquisition device does not capture an image frame at the detection time point, the electronic device can obtain an image frame captured by the image acquisition device most recently before the detection time point as the image frame captured at the detection time point.

[0211] The pixel position of the target of interest in the to-be-processed image frame represents an image area occupied by the target of interest in the to-be-processed image frame. For example, the pixel position of the target of interest in the to-be-processed image frame can be represented as: . is a pixel coordinate of a top-left pixel point of an image area occupied by the target of interest in the to-be-processed image frame, is a width of the image area occupied by the target of interest in the to-be-processed image frame, is a height of the image area occupied by the target of interest in the to-be-processed image frame.

[0212] The imaging plane represents a plane formed by an image coordinate system of a two-dimensional image (i.e., an image frame) formed by the lens in the image acquisition device when collecting an image. Moreover, the direction represented by the optical axis of the lens is perpendicular to the imaging plane.

[0213] The direction of the imaging plane includes two mutually perpendicular directions in the imaging plane. For ease of description, the width direction and the height direction of the image coordinate system can be determined as the two directions in the imaging plane (i.e., the direction of the imaging plane). The width direction of the image frame can be denoted as the direction represented by the x-axis, and the height direction of the image frame can be denoted as the direction represented by the y-axis.

[0214] In an implementation manner, the electronic device can obtain a pixel position (which can be referred to as a first historical pixel position) of the target of interest in an image frame collected by the image acquisition device at a detection time point preceding the first detection time point. Then, a pixel distance between the first historical pixel position and the pixel position of the target of interest in the to-be-processed image frame is calculated.

[0215] It can be understood that the rotation of the lens can cause a fixed position in the real world to have different pixel positions in image frames captured at different angles. Therefore, in order to determine the influence of the rotation of the lens on the change of the pixel position of the target of interest, the electronic device can obtain the rotation angle of the lens from the detection time point preceding the first detection time point to the first detection time point. Then, according to the field of view (FOV) of the lens, the amount of change of the pixel position caused by the rotation angle of the lens can be determined.

[0216] Then, the sum of the pixel distance between the first historical pixel position and the pixel position of the target of interest in the to-be-processed image frame and the amount of change of the pixel position caused by the rotation angle of the lens is obtained, and the ratio of the sum to the detection time interval is calculated as the moving speed (i.e., the first moving speed) of the target of interest in the direction of the imaging plane at the first detection time point.

[0217] The pixel distance between the first historical pixel position and the pixel position of the target of interest in the image frame to be processed can be: the distance between the pixel coordinates of the center point of the image region represented by the first historical pixel position and the pixel coordinates of the center point of the image region occupied by the target of interest in the image frame to be processed.

[0218] It is understandable that the obtained first moving velocity includes the velocity components of the target of interest in two directions within the imaging plane at the third detection time.

[0219] The movement speed of the target of interest at a given detection time can be denoted as: . This represents the velocity component of the target of interest in the direction represented by the x-axis at the detection moment; This represents the velocity component of the target of interest in the direction represented by the y-axis at the detection moment. The value of can be positive or negative. A positive value indicates that the target of interest has moved in the positive x-axis direction (to the right in the imaging plane) at the detection moment. A negative value indicates that the target of interest has moved in the opposite direction of the x-axis (to the left of the imaging plane) at the detection moment. Similarly, The value of can be positive or negative. A positive value indicates that the target of interest has moved in the positive y-direction (upper side of the imaging plane) at the detection moment. A negative value indicates that the target of interest moves in the opposite direction of the y-axis (the lower side of the imaging plane) at the detection moment.

[0220] In some embodiments, step S201 includes:

[0221] If the image quality of the image region occupied by the target of interest in the image frame meets the preset conditions, the moving speed of the target of interest in the direction of the imaging plane at the first detection time is calculated based on the pixel position of the target of interest in the image frame, and is used as the first moving speed.

[0222] The preset conditions include at least one of the following: the image area is not overexposed, the image area is not underexposed, and the image area is clear.

[0223] In this embodiment, after determining the pixel position of the target of interest in the image frame to be processed, the electronic device can perform image quality detection on the image region occupied by the target of interest in the image frame to be processed. The image region occupied by the target of interest in the image frame to be processed can also be referred to as the ROI (Region of Interest).

[0224] In one implementation, the image region occupied by the target of interest in the image frame to be processed is input into a pre-trained image quality detection model to obtain the image quality detection result of the image region occupied by the target of interest in the image frame to be processed.

[0225] The pre-trained image quality detection model can be a classification image model trained based on a deep learning network architecture. For example, the deep learning network architecture can be AlexNet (Alex Network), ResNet (Residual Neural Network), etc.

[0226] The image quality detection model is trained based on sample images and their corresponding labels. A label for a sample image indicates whether the image is overexposed, underexposed, or sharp.

[0227] In another implementation, the brightness of each pixel location in the image region occupied by the target of interest in each image frame to be processed can be calculated according to a preset pixel brightness calculation formula. Then, the average brightness of each pixel location in the image region occupied by the target of interest in each image frame to be processed is calculated, and it is determined whether the average value is within a specified brightness range. If the average value is within the specified brightness range, it indicates that the image region is neither overexposed nor underexposed. If the average value is greater than the specified brightness range, it indicates that the image region is overexposed. If the average value is less than the specified brightness range, it indicates that the image region is underexposed.

[0228] The brightness of a pixel can be represented as: . This represents the value of the red channel in the pixel value at that pixel location. This represents the value of the green channel in the pixel value at that pixel location. This represents the value of the blue channel in the pixel value at that pixel location.

[0229] In another implementation, the sharpness of the image region occupied by the target of interest in the image frame to be processed can be determined according to a preset sharpness evaluation function. For example, the sharpness evaluation function can be the Energy of Gradient (EOG) function.

[0230] Correspondingly, the electronic device can calculate the gray value of the pixel point in the image area occupied by the target of interest in the to-be-processed image frame. For each pixel point in the image area, the difference between the gray values of the pixel point and adjacent pixel points in the height direction and the width direction is calculated respectively. Then, for the height direction, the sum of the squares of the differences between the gray values of the pixel points in the height direction is calculated; for the width direction, the sum of the squares of the differences between the gray values of the pixel points in the width direction is calculated. The sum of the sums of the squares in the height direction and the width direction is calculated as the energy gradient (EOG) value. If the energy gradient (EOG) value is greater than a preset energy gradient threshold, it indicates that the image area is clear; if the energy gradient (EOG) value is not greater than the preset energy gradient threshold, it indicates that the image area is not clear.

[0231] The preset condition can be set by the technical personnel according to the actual needs. For example, in a scene with high requirements for image quality, the preset condition includes that the image area is not overexposed, the image area is not too dark, and the image area is clear.

[0232] Correspondingly, the first movement speed of the target of interest in the direction of the imaging plane at the first detection time can be calculated when the image quality of the image area occupied by the target of interest in the to-be-processed image frame meets the preset condition.

[0233] In an implementation manner, when the image quality of the image area occupied by the target of interest in the to-be-processed image frame does not meet the preset condition, the electronic device can not calculate the first movement speed and the movement speed of the target of interest in the direction of the imaging plane at the third detection time (i.e., the third movement speed). Correspondingly, when the second detection time is reached, the electronic device can maintain the current speed of the motor in the image acquisition device at the current time until the third detection time is reached.

[0234] Based on the above processing, the first movement speed can be calculated when the image quality of the image area occupied by the target of interest in the to-be-processed image frame meets the preset condition. It can be understood that if the image quality of the image area occupied by the target of interest in the to-be-processed image frame does not meet the preset condition, it indicates that the quality of the current to-be-processed image frame may not be high. Correspondingly, if the speed is predicted according to the to-be-processed image frame, the third movement speed of the target of interest at the third detection time obtained may not be accurate, which may interfere with the subsequent process of adjusting the speed of the motor. Therefore, only when the image quality of the image area occupied by the target of interest in the to-be-processed image frame meets the preset condition, the subsequent steps are executed, which can eliminate the influence of the image frame with low image quality on the subsequent control of the image acquisition device, and improve the accuracy of the control of the image acquisition device.

[0235] For step S202, the first detection time, the second detection time and the third detection time are three continuous adjacent detection times, wherein the second detection time is the next detection time of the first detection time, and the third detection time is the next detection time of the second detection time.

[0236] The historical detection time before the first detection time can be: the second specified number of continuous adjacent detection times before the first detection time. For example, the second specified number can be 5.

[0237] For each historical detection time before the first detection time, the process of obtaining the moving speed (i.e., the second moving speed) of the target of interest in the direction of the imaging plane at the historical detection time can refer to the process of calculating the first moving speed described above, which will not be repeated here.

[0238] Correspondingly, the third detection time moving speed of the target of interest in the direction of the imaging plane can be calculated according to the second moving speed and the first moving speed.

[0239] In some embodiments, referring to Figure 3 , Figure 3 A second flowchart for controlling after detecting the target of interest is provided for the embodiments of the present application. In Figure 2 , step S202 includes:

[0240] S2021: determining the pixel distance between the target of interest and the reference stationary object in the image frame.

[0241] S2022: inputting the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection time before the first detection time, and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection time as the third moving speed.

[0242] The speed prediction model is: trained based on sample data and corresponding sample labels.

[0243] The sample data includes: the pixel distance between the sample object and the reference stationary object in the sample image frame collected at a detection time, the moving speed of the sample object in the direction of the imaging plane at the historical detection time before the detection time, and the moving speed of the sample object in the direction of the imaging plane at the detection time; the sample label corresponding to one sample data includes: the speed of the sample object at the second detection time after the detection time.

[0244] In the embodiments of the present application, the electronic device can pre-acquire the pixel position of the reference stationary object in the to-be-processed image frame. The reference stationary object refers to an object that is stationary in the real world. For example, the reference stationary object can include trees, walls, roadblocks, and the like.

[0245] The process in which the electronic device acquires the pixel position of the reference stationary object in the to-be-processed image frame can refer to the process of acquiring the pixel position of the target of interest in the to-be-processed image frame in the above-described embodiments, and will not be described here.

[0246] Correspondingly, the electronic device can determine the pixel distance between the target of interest and the reference stationary object in the to-be-processed image frame according to the pixel position of the target of interest in the to-be-processed image frame and the pixel position of the reference stationary object in the to-be-processed image frame. For example, the distance (such as the Euclidean distance) between the center point pixel coordinates contained in the pixel position of the target of interest in the to-be-processed image frame and the center point pixel coordinates contained in the pixel position of the reference stationary object in the to-be-processed image frame can be calculated as the pixel distance between the target of interest and the reference stationary object in the to-be-processed image frame.

[0247] Further, the electronic device can use the pre-trained speed prediction model to predict the moving speed of the target of interest in the direction of the imaging plane at the third detection time (i.e., the third moving speed).

[0248] The pre-trained speed prediction model is obtained based on a time series network architecture. The time series network architecture can be a long short-term memory network (LSTM) structure, an auto-regression and moving average model (ARMA), or the like.

[0249] The speed prediction model is trained based on sample data and corresponding sample labels.

[0250] The sample data includes the pixel distance between the sample object and the reference stationary object in a sample image frame collected at a detection time, the moving speed of the sample object in the direction of the imaging plane at a historical detection time before the detection time, and the moving speed of the sample object in the direction of the imaging plane at the detection time. The sample label corresponding to one sample data includes the speed of the sample object at a second detection time after the detection time.

[0251] Further, the determined pixel distance, the second moving speed of the object of interest in the direction of the imaging plane at the historical detection moment before the first detection moment, and the first moving speed are input into a pre-trained speed prediction model to obtain the moving speed of the object of interest in the direction of the imaging plane at the third detection moment as the third moving speed.

[0252] The second moving speed and the first moving speed can reflect the change amount of the size and direction of the moving speed of the object of interest from the historical detection moment to the first detection moment, and accordingly, the speed prediction model can combine the change amount of the historical moving speed of the object of interest and the pixel distance between the object of interest and the reference stationary object in the to-be-processed image frame to predict the moving speed of the object of interest in the direction of the imaging plane at the third detection moment.

[0253] Based on the above processing, in the process of predicting the moving speed of the object of interest in the direction of the imaging plane at the third detection moment, the electronic device can combine the historical and current moving speeds of the object of interest, predict based on the pre-trained speed prediction model, and introduce the distance between the object of interest and the reference stationary object in the to-be-processed image frame in the prediction process. It can be understood that since the lens rotation will cause the fixed position in the real world to present different pixel positions in the image frames under different shooting angles, the introduction of the reference stationary object can construct the position correlation relationship between the stationary object and the moving object of interest in the real world, thereby providing more abundant and stable information for predicting the moving speed of the object of interest, effectively reducing the interference caused by the lens rotation, and enhancing the accuracy of the speed prediction.

[0254] In another implementation manner, the electronic device can calculate the moving speed of the object of interest in the direction of the imaging plane at the third detection moment as the third moving speed based on a preset speed prediction algorithm according to the second moving speed of the object of interest in the direction of the imaging plane at the historical detection moment before the first detection moment and the first moving speed. The preset speed prediction algorithm can be a Kalman filter algorithm.

[0255] In another implementation manner, the electronic device can calculate the moving speed of the object of interest in the direction of the imaging plane at the third detection moment as the third moving speed based on a preset speed prediction algorithm according to the second moving speed of the object of interest in the direction of the imaging plane at the historical detection moment before the first detection moment and the first moving speed. The preset speed prediction algorithm can be a Kalman filter algorithm.

[0256] The speed prediction model is obtained by training based on sample data and corresponding sample labels. The sample data includes: a moving speed of the sample object in the direction of the imaging plane at a historical detection time before a detection time, and a moving speed of the sample object in the direction of the imaging plane at the detection time; and the sample label corresponding to the sample data includes: a speed of the sample object at a second detection time after the detection time.

[0257] The pre-trained speed prediction model is obtained by training based on a time series network architecture. The time series network architecture can be a long short term memory network (LSTM) structure, an auto-regression and moving average model (ARMA), or the like.

[0258] Based on the above processing, the electronic device can use the speed prediction model obtained by training based on the time series network architecture to predict the change trend of the moving speed of the target of interest in the historical time period, and obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection time.

[0259] For step S203, when the second detection time is reached, the electronic device can adjust the rotation speed of the motor according to the third moving speed, the current rotation speed of the motor in the image acquisition device, and the maximum rotation speed supported by the motor.

[0260] In some embodiments, referring to Figure 4 , Figure 4 A third flowchart for controlling after detecting the target of interest is provided for the embodiments of the present application. In Figure 2 , step S203 includes:

[0261] S2031: When the second detection time is reached, a peak rotation speed required by the motor at a specified time between the second detection time and the third detection time is determined according to the third moving speed, the current rotation speed of the motor in the image acquisition device, and the maximum rotation speed supported by the motor.

[0262] S2032: The rotation speed of the motor is adjusted in an incremental or decremental manner with the current rotation speed of the motor as the initial rotation speed, so that the rotation speed of the motor reaches the peak rotation speed at the specified time.

[0263] S2033: When the specified time is reached, if the current rotation speed of the motor does not reach the maximum rotation speed supported, the rotation speed of the motor is adjusted in an incremental or decremental manner with the current rotation speed of the motor as the initial rotation speed, so that the rotation speed of the motor reaches the third moving speed at the third detection time.

[0264] The method further includes:

[0265] S204: When the specified time is reached, if the current speed of the motor at the specified time is the maximum speed supported, the motor is controlled to maintain the current speed until the third detection time.

[0266] In the embodiments of the present application, the specified time between the second detection time and the third detection time is the time corresponding to the specified duration after the second detection time. For example, the specified duration is a preset multiple of a detection time interval. The preset multiple can be in the range of [0.5, 0.8]. For example, the preset multiple can be 2 / 3.

[0267] When the second detection time is reached, the electronic device can determine the peak speed required for the motor to reach at the specified time between the second detection time and the third detection time.

[0268] The electronic device can adjust the speed of the motor in the image acquisition device to adjust the angle (i.e. the direction of the optical axis) of the lens in the image acquisition device to capture different areas in the real world.

[0269] At any time, the speed of the motor represents the size and direction of the angle of the motor at that time.

[0270] For example, the speed of the motor at any time can be represented as: . represents the speed component of the motor in the direction represented by the x-axis in the imaging plane; represents the speed component of the motor in the direction represented by the y-axis in the imaging plane. Wherein, The value of can be positive and negative, The value of is positive, indicating that the motor rotates in the positive direction of the x-axis (right side of the imaging plane), The value of is negative, indicating that the motor rotates in the negative direction of the x-axis (left side of the imaging plane). Similarly, The value of can be positive and negative, The value of is positive, indicating that the motor rotates in the positive direction of the y-axis (upper side of the imaging plane), The value of is negative, indicating that the motor rotates in the negative direction of the y-axis (lower side of the imaging plane).

[0271] It can be understood that as the speed of the motor increases, the quality of the image frames captured by the image acquisition device will generally decrease. Therefore, the skilled person can determine the maximum speed supported by the motor based on the relationship between the speed of the motor and the image quality of the captured image frames, to ensure that the image quality of the captured image frames meets the preset conditions when the motor rotates at the maximum speed.

[0272] For step S2032, after determining the peak speed of the motor at the specified time, the electronic device can adjust the speed of the motor in an increasing or decreasing manner, taking the current speed of the motor as the initial speed, so that the speed of the motor reaches the peak speed at the specified time.

[0273] The electronic device adjusts the speed of the motor by adjusting the pulse frequency. The higher the pulse frequency, the faster the motor rotates; the lower the pulse frequency, the slower the motor rotates.

[0274] In one implementation, the electronic device can calculate the speed difference between the current speed of the motor and the peak speed of the motor at the specified time, and calculate the time difference between the current time and the specified time. Then, the ratio of the speed difference and the time difference is calculated to obtain the speed acceleration of the motor from the current time to the specified time. Accordingly, the pulse frequency is adjusted in a linear adjustment manner according to the corresponding relationship between the pulse frequency and the speed of the motor, so that the motor maintains the speed acceleration represented by the ratio, thereby ensuring that the motor reaches the peak speed at the specified time.

[0275] In another implementation, the electronic device can calculate the speed difference between the current speed of the motor and the peak speed of the motor at the specified time, and calculate the time difference between the current time and the specified time. Then, the calculated speed difference and time difference are fitted using a sine function, so that the electronic device adjusts the speed of the motor in a stepwise manner to make the speed of the motor approach the sine function. Accordingly, the pulse frequency is adjusted according to the corresponding relationship between the pulse frequency and the speed of the motor, so that the motor accelerates or decelerates according to the fitting result, ensuring that the motor reaches the peak speed at the specified time. In addition, the speed of the motor at each time from the current time to the specified time is determined by fitting the sine function, which can gradually adjust the speed of the motor and ensure the smoothness of the motor rotation.

[0276] When the specified time is reached, if the current speed of the motor does not reach the maximum speed supported, the speed of the motor is adjusted in an increasing or decreasing manner from the specified time, taking the current speed of the motor as the initial speed, so that the speed of the motor reaches the third moving speed at the third detection time.

[0277] For example, when the specified time is reached, if the current speed of the motor is consistent with the direction represented by the third moving speed, and the magnitude of the current speed of the motor is greater than the magnitude of the third moving speed, the speed of the motor is adjusted in a decreasing manner, taking the current speed of the motor as the initial speed.

[0278] If the rotation speed of the motor at the current time is consistent with the direction represented by the third moving speed, and the magnitude of the rotation speed of the motor at the current time is less than the magnitude of the third moving speed, the rotation speed of the motor is adjusted in an increasing manner with the rotation speed of the motor at the current time as the initial rotation speed.

[0279] If the rotation speed of the motor at the current time is opposite to the direction represented by the third moving speed, the magnitude of the rotation speed of the motor is adjusted to 0 in a decreasing manner with the rotation speed of the motor at the current time as the initial rotation speed, and then the rotation speed of the motor is adjusted in an increasing manner with 0 as the initial rotation speed.

[0280] Correspondingly, the rotation speed of the lens in the image acquisition device at the third detection time is consistent with the third moving speed, and the lens rotation speed is adjusted starting from the second detection time, so that the pixel position change speed of the target of interest in the image frame caused by the rotation of the lens in the image acquisition device at the third detection time is consistent with the pixel position change speed of the target of interest in the image frame caused by the movement of the target of interest itself, which can ensure the synchronization of the rotation speed of the lens in the image acquisition device and the moving speed of the target of interest, that is, the lens rotation speed can be adjusted according to the moving speed of the target of interest to enable the lens in the image acquisition device to continuously capture the moving target of interest.

[0281] And from the second detection time to the third detection time, the movement distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens, that is, within the time interval from the second detection time to the third detection time, the pixel position of the target of interest in the image caused by the movement of the target of interest itself can be compensated by the pixel position of the target of interest in the image caused by the rotation of the lens, to ensure the synchronization of the rotation amount of the lens in the image acquisition device and the movement amount of the target of interest, that is, the movement amount of the target of interest is compensated by the rotation amount of the lens, so that the position of the target of interest in the image frame collected within the time interval is stable, and the target of interest is kept in the middle position of the image frame collected by the image acquisition device to a certain extent.

[0282] When the specified time is reached, if the rotation speed of the motor at the current time is the maximum supported rotation speed, it indicates that the motor may still be unable to enable the image acquisition device to continuously capture the target of interest when rotating at the maximum supported rotation speed. At this time, the motor can be controlled to maintain the rotation speed at the current time until the third detection time to ensure that the image acquisition device follows the target of interest to the maximum extent.

[0283] It can be understood that if the first detection time is not the first detection time in the continuous shooting process of the image acquisition device on the target of interest, at the last detection time of the first detection time, the electronic device can perform speed prediction according to the above steps S201 to S202 to obtain the moving speed of the target of interest in the direction of the imaging plane at the second detection time. When the first detection time is reached, the speed of the motor is adjusted. And the way of adjusting the speed of the motor from the first detection time to the second detection time can refer to the way of adjusting the speed of the motor from the second detection time to the third detection time described above, which will not be repeated here.

[0284] In some embodiments, S2031, comprising:

[0285] When the second detection time is reached, according to the third moving speed, the current speed of the motor in the image acquisition device at the moment, and the maximum speed supported by the motor, the peak speed required by the motor at the specified time between the second detection time and the third detection time is determined according to the preset formula.

[0286] Wherein, the preset formula is represented as:

[0287]

[0288] The peak speed is represented as: The maximum speed supported by the motor is represented as: The third moving speed is represented as: The speed of the motor at the second detection time is represented as:

[0289] In the embodiments of the present application, The preset formula is set by the technician in advance, such as: The preset formula is set by the technician in advance, such as:

[0290] It can be understood that the peak speed required by the motor at the specified time includes: the peak speed in the x-axis direction of the imaging plane, and the peak speed in the y-axis direction of the imaging plane.

[0291] The maximum speed supported by the motor includes: the maximum speed supported by the motor in each direction of the imaging plane. Generally, the maximum speed supported by the motor in each direction of the imaging plane is consistent.

[0292] Based on the above processing, it is ensured that the peak speed required by the motor at the specified time between the second detection time and the third detection time is within the maximum speed supported by the motor. In this way, it can be ensured that the image quality of the image frame collected by the image acquisition device in the case of rotating at the peak speed, and the stability of the image acquisition device control is ensured.

[0293] Taking the determination of the peak rotational speed in the x-axis direction within the imaging plane as an example, and in The value is 1.5, and the maximum speed supported by the motor is 7, meaning the motor supports speeds in the range of [-7, 7].

[0294] If the motor's rotational speed is 3 at the second detection moment and its third moving speed is 6, then according to the preset formula, it can be known that... The value is 7.5, which is greater than the maximum speed supported by the motor, 7. Therefore, the peak speed that the motor needs to reach at the specified time is 7. Accordingly, when the second detection time is reached, the electronic equipment can control the motor to adjust the motor speed from 3 in an incremental manner so that the motor speed reaches 7 at the specified time, and control the motor to maintain the current speed of 7 until the third detection time.

[0295] If the motor's speed is 6 at the second detection moment and its moving speed is 3 at the third moving speed, then according to the preset formula, The value is 1.5, which is no greater than the maximum speed supported by the motor (7). Therefore, the peak speed that the motor needs to reach at the specified time is 1.5. Accordingly, when the second detection time is reached, the electronic equipment can control the motor to adjust its speed from 6 in a decreasing manner so that the motor speed reaches 1.5 at the specified time. When the specified time is reached, the electronic equipment can control the motor to adjust its speed from 1.5 in an increasing manner so that the motor speed reaches 3 at the third detection time.

[0296] If the motor's rotational speed at the second detection moment is -3 and its third moving speed is 3, then according to the preset formula, it can be known that... The value is 6, which is no greater than the maximum speed supported by the motor, 7. Therefore, the peak speed that the motor needs to reach at the specified time is 6. Accordingly, when the second detection time is reached, the electronic equipment can control the motor to adjust its speed from -3 in a decreasing manner to 0; then, starting from 0, the motor speed is adjusted in an increasing manner so that the motor speed reaches 6 at the specified time. When the specified time is reached, the electronic equipment can control the motor to adjust its speed from 6 in a decreasing manner so that the motor speed reaches 3 at the third detection time.

[0297] If the motor's rotational speed is 3 at the second detection moment and its third moving speed is -3, then according to the preset formula, we can know that... The value of the peak speed is -6, which is not greater than the maximum speed -7 supported by the motor. Therefore, the peak speed required by the motor at the specified moment is -6. Correspondingly, when the second detection moment is reached, the electronic device can control the motor to adjust the speed of the motor in a decreasing manner from 3 to 0, and then adjust the speed of the motor in an increasing manner from 0, so that the speed of the motor reaches -6 at the specified moment. When the specified moment is reached, the electronic device can control the motor to adjust the speed of the motor in a decreasing manner from -6, so that the speed of the motor reaches -3 at the third detection moment.

[0298] The process of determining the peak speed in the y-axis direction of the imaging plane can refer to the process of determining the peak speed in the x-axis direction of the imaging plane, which will not be repeated here.

[0299] Referring to Figure 5 , Figure 5 The flowchart diagram of the image acquisition device control provided by the embodiment of the present application is shown. Figure 5 In the embodiment of the present application,

[0300] Target detection: When the first detection moment is reached, the image frame captured by the image acquisition device is obtained, and the pixel positions of each object in the image frame are obtained.

[0301] Quality judgment: The electronic device can detect the image quality of the image frame and / or the image quality of the image area occupied by each object. When the image quality of the image frame meets the preset condition, and / or the image quality of the image area occupied by each object meets the preset condition, subsequent processing is performed.

[0302] Target following: According to the pixel positions of each object contained in the image frame and the pixel positions of the target of interest in the captured historical image frame, the pixel position of the target of interest in the image frame is determined.

[0303] Speed prediction: According to the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment and the moving speed of the target of interest in the direction of the imaging plane at the first detection moment (i.e. the first moving speed), the moving speed of the target of interest in the direction of the imaging plane at the third detection moment is calculated as the third moving speed.

[0304] The motor parameter adjustment indicates that when the second detection moment is reached, the rotation speed of the motor is adjusted according to the third movement speed, the current rotation speed of the motor in the image acquisition device at the moment, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third movement speed, and the movement distance of the target of interest in the direction of the imaging plane from the second detection moment to the third detection moment is consistent with the rotation angle of the lens.

[0305] It should be noted that the sample image and the sample data in the embodiment are both from a public data set.

[0306] Based on the same inventive concept, an embodiment of the present application provides an intelligent image processing device for open self-defined detection targets. Referring to Figure 6 , Figure 6 A structural schematic diagram of an intelligent image processing device provided by an embodiment of the present application.

[0307] The intelligent image processing device 600 comprises:

[0308] The lens 601 is configured to acquire image frames.

[0309] The processor 602 is configured to perform the intelligent image processing method for open self-defined detection targets.

[0310] Based on the intelligent image processing device for open self-defined detection targets provided by the embodiment of the present application, the object description information for determining the target of interest of the user can be input in a self-defined input manner. Correspondingly, whether the target of interest exists in the acquired image frames can be detected based on the object description information. Furthermore, if the target of interest exists in the image frames, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frames. That is, the intelligent image processing device for open self-defined detection targets can continuously capture the target (i.e., the target of interest) in the image frames that meets the object description information, and has high usability.

[0311] In some embodiments, the device further comprises a gimbal.

[0312] The processor is specifically configured to send a control instruction to the gimbal.

[0313] The gimbal is configured to adjust the image acquisition direction of the lens according to the control instruction sent by the processor, so that the target of interest is located at a preset position in the image frames.

[0314] In the embodiment of the present application, the intelligent image processing device includes a holder, and a motor for adjusting the image capturing direction of the lens is arranged in the holder. Correspondingly, the holder can control the motor to rotate according to the control instruction sent by the processor, so as to adjust the image capturing direction of the lens, and make the target of interest located at the preset position in the image frame.

[0315] Based on the same inventive concept, the embodiment of the present application provides an open intelligent image processing system for customizing detection target. Figure 7 , Figure 7 FIG. 1 is a structural schematic diagram of an intelligent image processing system provided by the embodiment of the present application.

[0316] The intelligent image processing system 700 includes:

[0317] The client 701 is configured to receive the object description information input by the user in a customized manner.

[0318] The customized input manner includes at least one of the following: inputting text, inputting image, inputting audio, and triggering an interactive instruction containing position information in the image frame; and the object description information is used to determine the target of interest of the user.

[0319] The image processing end 702 is configured to capture an image frame, and detect whether the target of interest exists in the captured image frame based on the object description information; and if the target of interest exists in the image frame, control the direction of image capturing, so as to make the target of interest located at the preset position in the image frame.

[0320] Based on the open intelligent image processing system for customizing detection target provided by the embodiment of the present application, the client supports the user to input the object description information used to determine the target of interest of the user in a customized manner. Correspondingly, the image processing end can detect whether the target of interest exists in the captured image frame based on the object description information. Further, if the target of interest exists in the image frame, the direction of image capturing is controlled, so as to make the target of interest located at the preset position in the image frame. That is, the electronic device can control the image capturing device to continuously capture the target (i.e., the target of interest) conforming to the object description information in the image frame according to the actual demand of the user, and improve the usability of the image capturing device.

[0321] In the embodiment of the present application, the client can provide an interactive interface for receiving the object description information input by the user, and the client can communicate with the image processing end. Correspondingly, the user can input the object description information through the interactive interface of the client, and then the client can send the object description information input by the user to the image processing end.

[0322] In an implementation, the image processing end can be an intelligent image processing device that supports open self-defined detection targets as described in the above embodiments.

[0323] In some embodiments, the client is integrated in the image processing end.

[0324] In the embodiments of the present application, the client can be integrated in the image processing end. Accordingly, the image processing end can provide an interactive interface for receiving object description information input by a user.

[0325] In some embodiments, the system further includes a cloud platform.

[0326] The cloud platform is configured to receive the object description information sent by the client and forward the object description information to the image processing end.

[0327] In the embodiments of the present application, the client can communicate with the cloud platform. Accordingly, the client can send the object description information input by a user to the cloud platform. Further, the cloud platform can forward the object description information to the image processing end.

[0328] In the case where the image processing end has high performance, the image processing end can perform the above steps S102 to S103. In this way, the computing and communication pressure of the cloud platform can be dispersed, and the efficiency of intelligent image processing can be improved.

[0329] Based on the same inventive concept, the embodiments of the present application provide an intelligent image processing system that supports open self-defined detection targets. Referring to Figure 8 , Figure 8 Another structure diagram of an intelligent image processing system is provided for the embodiments of the present application.

[0330] The intelligent image processing system 800 includes:

[0331] The client 801 is configured to receive object description information input by a user in a self-defined manner. The self-defined manner includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame. The object description information is used to determine a target of interest of the user.

[0332] The image processing end 802 is configured to collect image frames.

[0333] The cloud platform 803 is configured to receive the object description information sent by the client and receive the image frames sent by the image processing end. Based on the object description information, the cloud platform 803 detects whether the target of interest exists in the image frames and sends a detection result to the image processing end.

[0334] The image processing end 802 is further configured to control a direction of image acquisition according to the detection result sent by the cloud platform, so that the target of interest is located at a preset position in the image frame.

[0335] Based on the open self-defined target detection intelligent image processing system provided in the embodiments of the present application, the client supports the user to input object description information for determining the target of interest of the user in a self-defined input manner. Correspondingly, the cloud platform can detect whether the target of interest exists in the collected image frame based on the object description information. Further, if the target of interest exists in the image frame, the image processing end controls the direction of image acquisition, so that the target of interest is located at a preset position in the image frame. That is, the electronic device can control the image acquisition device to continuously capture the target (i.e., the target of interest) conforming to the object description information in the image frame according to the actual needs of the user, thereby improving the usability of the image acquisition device.

[0336] In the embodiments of the present application, the image processing end includes a lens and a motor for controlling the direction of image acquisition.

[0337] In the case that the performance of the cloud platform is high, the cloud platform can perform the above steps S102 to S103. In this way, the application range of the intelligent image processing method can be ensured, and continuous capturing of the target of interest can be realized in the case that the performance of the image processing end is limited.

[0338] Based on the same inventive concept, the embodiments of the present application provide an open self-defined target detection intelligent image processing device. Referring to Figure 9 , Figure 9 The structure diagram of an open self-defined target detection intelligent image processing device provided in the embodiments of the present application is shown in FIG. 9. The device includes:

[0339] The description information acquisition module 901 is configured to acquire object description information input by the user in a self-defined manner. The self-defined input manner includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame. The object description information is used to determine the target of interest of the user.

[0340] The detection module 902 is configured to detect whether the target of interest exists in the collected image frame based on the object description information.

[0341] The first control module 903 is configured to control the direction of image acquisition, so that the target of interest is located at a preset position in the image frame, if it is detected that the target of interest exists in the image frame.

[0342] In some embodiments, the data format of the object description information is one of text, an image, and audio.

[0343] The detection module 902 comprises:

[0344] The first detection submodule is configured to, for a collected image frame, if it is detected that the object of interest does not exist in a continuous specified number of image frames before the image frame, or the image frame is the first image frame collected after the object description information is received, detect whether the object of interest exists in the image frame based on the object description information and the image frame and a pre-trained open-vocabulary object detection model.

[0345] The second detection submodule is configured to, if it is detected that the object of interest exists in at least one of the continuous specified number of image frames, detect the image frame by using the open-vocabulary object detection model, and determine whether the object of interest exists in the image frame according to a pixel position of a detected target and pixel positions of the objects of interest detected in the continuous specified number of image frames.

[0346] In some embodiments, the second detection submodule is specifically configured to:

[0347] input the object description information and the image frame into a pre-trained open-vocabulary object detection model to obtain a pixel position of a target in the image frame that meets the object description information;

[0348] or,

[0349] input the image frame and a preset prompt word into a pre-trained open-vocabulary object detection model to obtain a pixel position of a target in the image frame that meets the preset prompt word.

[0350] In some embodiments, the object description information is triggered by a user in a specified image frame.

[0351] The detection module 902 comprises:

[0352] The third detection submodule is configured to obtain a pixel position of each target in the specified image frame by using a pre-trained open-vocabulary object detection model, and determine whether the object of interest exists in the image frame according to the determined pixel position of each target and position information contained in the object description information.

[0353] a fourth detection submodule configured to, for each image frame after the specified image frame, if it is detected that there is at least one image frame of the continuous specified number of image frames before the image frame that has the target of interest, obtain pixel positions of each target in the image frame by using the open-vocabulary target detection model, and determine whether the target of interest exists in the image frame according to the determined pixel positions of each target and the pixel positions of the target of interest detected in the continuous specified number of image frames.

[0354] In some embodiments, the first control module 903 includes:

[0355] a first speed calculation submodule configured to, if the target of interest exists in the image frame obtained at the first detection time, calculate a moving speed of the target of interest in the direction of the imaging plane at the first detection time as a first moving speed based on the pixel positions of the target of interest in the image frame.

[0356] a second speed calculation submodule configured to calculate a moving speed of the target of interest in the direction of the imaging plane at a third detection time as a third moving speed according to a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time and the first moving speed, wherein the third detection time is a next detection time of the second detection time, and the second detection time is a next detection time of the first detection time.

[0357] a control submodule configured to, when the second detection time is reached, adjust a rotating speed of a motor in the image acquisition device according to the third moving speed, a current rotating speed of the motor at the current time, and a maximum rotating speed supported by the motor, so that a rotating speed of a lens in the image acquisition device at the third detection time is consistent with the third moving speed, and a moving distance of the target of interest in the direction of the imaging plane from the second detection time to the third detection time is consistent with a rotating angle of the lens.

[0358] In some embodiments, the second speed calculation submodule is specifically configured to:

[0359] determine a pixel distance between the target of interest and a reference stationary object in the image frame;

[0360] input the determined pixel distance, a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time, and the first moving speed into a pre-trained speed prediction model to obtain a moving speed of the target of interest in the direction of the imaging plane at the third detection time as a third moving speed.

[0361] The speed prediction model is obtained based on sample data and corresponding sample labels.

[0362] The sample data includes a pixel distance between a sample object and a reference stationary object in a sample image frame collected at a detection time, a moving speed of the sample object in a direction of an imaging plane at a historical detection time before the detection time, and a moving speed of the sample object in the direction of the imaging plane at the detection time. The sample label corresponding to the sample data includes a speed of the sample object at a second detection time after the detection time.

[0363] In some embodiments, the first speed calculation sub-module is specifically configured to:

[0364] In a case where an image quality of an image region occupied by the target of interest in the image frame meets a preset condition, the moving speed of the target of interest in the direction of the imaging plane at the first detection time is calculated as the first moving speed based on a pixel position of the target of interest in the image frame.

[0365] The preset condition includes at least one of the following: the image region is not overexposed, the image region is not too dark, and the image region is clear.

[0366] In some embodiments, the control sub-module is specifically configured to:

[0367] The peak speed determination unit is configured to, when the second detection time is reached, determine a peak speed required by the motor at a specified time between the second detection time and the third detection time according to the third moving speed, a current speed of the motor at the current time, and a maximum speed supported by the motor.

[0368] The first speed adjustment unit is configured to take the current speed of the motor at the current time as an initial speed, and adjust the speed of the motor in an increasing or decreasing manner, so that the speed of the motor reaches the peak speed at the specified time.

[0369] The second speed adjustment unit is configured to, when the specified time is reached, if the current speed of the motor at the current time does not reach the maximum speed supported, take the current speed of the motor at the current time as an initial speed, and adjust the speed of the motor in an increasing or decreasing manner, so that the speed of the motor reaches the third moving speed at the third detection time.

[0370] The device further includes:

[0371] The second control module is configured to, when the specified time is reached, if the current rotating speed of the motor is the maximum rotating speed supported by the motor, control the motor to keep the current rotating speed until the third detection time.

[0372] In some embodiments, the peak rotating speed determination unit is specifically configured to:

[0373] When the second detection time is reached, according to a preset formula, the third moving speed, the current rotating speed of the motor at the second detection time, and the maximum rotating speed supported by the motor, a peak rotating speed required by the motor to reach at a specified time between the second detection time and the third detection time is determined.

[0374] The preset formula is represented as:

[0375]

[0376] The peak rotating speed is represented as ωp; The maximum rotating speed supported by the motor is represented as ωmax; The third moving speed is represented as v3; The rotating speed of the motor at the second detection time is represented as ω2.

[0377] Embodiments of the present application also provide an electronic device, as shown in the accompanying drawings, comprising: Figure 10

[0378] The memory 1001 is configured to store a computer program.

[0379] The processor 1002 is configured to execute the program stored in the memory 1001, and implement the steps of the open self-defined detection target intelligent image processing method.

[0380] The electronic device can further comprise a communication bus and / or a communication interface. The processor 1002, the communication interface, and the memory 1001 can communicate with each other through the communication bus.

[0381] The communication bus of the electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, and a control bus. For the convenience of representation, only one thick line is shown in the figure, but it does not mean that there is only one bus or only one type of bus.

[0382] The communication interface is configured to communicate between the electronic device and other devices.​

[0383] The memory can include a Random Access Memory (RAM) and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0384] The processor described above can be a general processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0385] In another embodiment provided in the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps of the open self-defined detection target intelligent image processing method in any of the above embodiments.

[0386] In another embodiment provided in the present application, a computer program product containing instructions is also provided, and when the computer program product is run on a computer, the computer is caused to execute the open self-defined detection target intelligent image processing method in any of the above embodiments.

[0387] In the embodiments described above, all or some of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or some of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded into and executed by a computer, all or some of the processes or functions according to the embodiments described in the specification are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a solid state disk (SSD) and the like.

[0388] It should be noted that, in this document, the terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0389] Each of the embodiments in the specification is described in a related manner, and the same or similar parts between each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for system, device, electronic device, computer readable storage medium embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0390] The above merely provides the preferred embodiment of the present application, and not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An open self-defined detection target intelligent image processing method, characterized in that, The method comprises: acquiring object description information input by a user; wherein the input mode comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; the object description information is used to determine a target of interest of the user; based on the object description information, detecting whether the target of interest exists in a collected image frame; if the target of interest is detected in the image frame, controlling the direction of image acquisition so that the target of interest is located at a preset position in the image frame; if the target of interest is detected in the image frame, controlling the direction of image acquisition so that the target of interest is located at a preset position in the image frame, comprising: if the target of interest exists in an image frame acquired at a first detection time, calculating the moving speed of the target of interest in the direction of the imaging plane at the first detection time based on the pixel position of the target of interest in the image frame, as a first moving speed; calculating the moving speed of the target of interest in the direction of the imaging plane at a third detection time as a third moving speed according to a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time and the first moving speed; wherein the third detection time is the next detection time of a second detection time; the second detection time is the next detection time of the first detection time; the time interval between every two adjacent detection times is preset according to the execution time required for executing a speed prediction process once; when the second detection time is reached, determining the peak speed required by the motor to reach at a specified time between the second detection time and the third detection time according to the third moving speed, the current speed of the motor at the current time, and the maximum speed supported by the motor; taking the current speed of the motor at the current time as an initial speed, and adjusting the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the peak speed at the specified time; when the specified time is reached, if the current speed of the motor at the current time does not reach the maximum speed supported, taking the current speed of the motor at the current time as an initial speed, and adjusting the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the third moving speed at the third detection time; the method further comprises: when the specified time is reached, if the current speed of the motor at the current time is the maximum speed supported, controlling the motor to maintain the current speed until the third detection time.

2. The method of claim 1, wherein, The data format of the object description information is one of text, image and audio; the detection of whether the target of interest exists in the collected image frame based on the object description information comprises: For a collected image frame, if it is detected that the target of interest does not exist in a continuous specified number of image frames before the image frame, or the image frame is the first image frame collected after receiving the object description information, then based on the object description information and the image frame, a pre-trained open-vocabulary target detection model is used to detect whether the target of interest exists in the image frame; If it is detected that the target of interest exists in at least one of the continuous specified number of image frames, then the open-vocabulary target detection model is used to detect the image frame; and whether the target of interest exists in the image frame is determined according to the pixel position of the detected target and the pixel position of the target of interest detected in the continuous specified number of image frames.

3. The method of claim 2, wherein, The detection of the image frame using the open-vocabulary target detection model includes: The object description information and the image frame are input into a pre-trained open-vocabulary target detection model to obtain the pixel position of the target in the image frame that meets the object description information; Or, The image frame and a preset prompt word are input into a pre-trained open-vocabulary target detection model to obtain the pixel position of the target in the image frame that meets the preset prompt word.

4. The method of claim 1, wherein, The object description information is triggered by a user in a specified image frame; The detection of whether the target of interest exists in the collected image frame based on the object description information includes: The pixel position of each target in the specified image frame is obtained using a pre-trained open-vocabulary target detection model, and whether the target of interest exists in the image frame is determined according to the determined pixel position of each target and the position information contained in the object description information; For each image frame after the specified image frame, if it is detected that the target of interest exists in at least one of the continuous specified number of image frames before the image frame, then the pixel position of each target in the image frame is obtained using the open-vocabulary target detection model, and whether the target of interest exists in the image frame is determined according to the determined pixel position of each target and the pixel position of the target of interest detected in the continuous specified number of image frames.

5. The method of claim 1, wherein, The calculation of the moving speed of the target of interest in the direction of the imaging plane at the third detection time as the third moving speed based on the second moving speed of the target of interest in the direction of the imaging plane at the historical detection time before the first detection time and the first moving speed includes: Determining the pixel distance between the target of interest and a reference stationary object in the image frame; Inputting the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection time before the first detection time, and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection time as the third moving speed; The speed prediction model is trained based on sample data and corresponding sample labels. The sample data comprises: a pixel distance between a sample object and a reference stationary object in a sample image frame collected at a detection time, a moving speed of the sample object in a direction of an imaging plane at a historical detection time before the detection time, and a moving speed of the sample object in the direction of the imaging plane at the detection time; and a sample label corresponding to the sample data comprises: a speed of the sample object at a second detection time after the detection time.

6. The method of claim 1, wherein, The first moving speed of the object of interest in the direction of the imaging plane at the first detection time is calculated based on the pixel position of the object of interest in the image frame, and the first moving speed comprises: In a case where an image quality of an image region occupied by the object of interest in the image frame meets a preset condition, the first moving speed of the object of interest in the direction of the imaging plane at the first detection time is calculated based on the pixel position of the object of interest in the image frame. The preset condition comprises at least one of the following: the image region is not overexposed, the image region is not too dark, and the image region is clear.

7. The method of claim 1, wherein, When the second detection time is reached, a peak speed required by the motor at a specified time between the second detection time and the third detection time is determined according to the third moving speed, a current speed of the motor at the current time, and a maximum speed supported by the motor. When the second detection time is reached, a peak speed required by the motor at a specified time between the second detection time and the third detection time is determined according to the third moving speed, a current speed of the motor at the current time, and a maximum speed supported by the motor, according to a preset formula. The preset formula is represented as: ; represents the peak rotational speed; represents the maximum rotational speed supported by the electric motor; represents the third movement speed; represents the rotational speed of the electric motor at the second detection instant.

8. An open self-defined detection target intelligent image processing device, characterized in that, The device comprises: a lens for collecting an image frame; a processor for executing the method of any one of claims 1-7.

9. The apparatus of claim 8, wherein, The device further comprises a gimbal. The processor is specifically configured to send a control instruction to the gimbal. The gimbal is configured to adjust an image collection direction of the lens according to the control instruction sent by the processor, so that the object of interest is located at a preset position in the image frame.

10. An open self-defined detection target intelligent image processing system, characterized in that, The system comprises: a client configured to receive object description information input by a user in a customized manner; the customized input manner comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; and the object description information is used to determine an object of interest of the user; an image processing end configured to collect an image frame, and detect whether the object of interest exists in the collected image frame based on the object description information; if the object of interest exists in the image frame, the direction of image collection is controlled so that the object of interest is located at a preset position in the image frame; the image processing end is specifically configured to: If the image frame obtained at the first detection time contains the target of interest, a moving speed of the target of interest in the imaging plane direction at the first detection time is calculated as a first moving speed based on a pixel position of the target of interest in the image frame; A third moving speed of the target of interest in the imaging plane direction at a third detection time is calculated as a third moving speed according to a second moving speed of the target of interest in the imaging plane direction at a historical detection time before the first detection time and the first moving speed, wherein the third detection time is a next detection time of a second detection time, the second detection time is a next detection time of the first detection time, and a time interval between every two adjacent detection times is preset according to an execution time required for performing a speed prediction process once; When the second detection time is reached, a peak speed required for the motor to reach at a specified time between the second detection time and the third detection time is determined according to the third moving speed, a current speed of the motor at the current time, and a maximum speed supported by the motor; The speed of the motor at the current time is taken as an initial speed, and the speed of the motor is adjusted in an increasing or decreasing manner so that the speed of the motor reaches the peak speed at the specified time; When the specified time is reached, if the speed of the motor at the current time does not reach the maximum speed supported, the speed of the motor at the current time is taken as an initial speed, and the speed of the motor is adjusted in an increasing or decreasing manner so that the speed of the motor reaches the third moving speed at the third detection time; The image processing end is further configured to: When the specified time is reached, if the speed of the motor at the current time is the maximum speed supported, the motor is controlled to maintain the speed at the current time until the third detection time.

11. The system of claim 10, wherein, The client is integrated in the image processing end.

12. The system of claim 10, wherein, The system further comprises a cloud platform. The cloud platform is configured to receive the object description information sent by the client and forward the object description information to the image processing end.

13. An open self-defined detection target intelligent image processing system, characterized in that, The system comprises: A client configured to receive object description information input by a user in a customized manner; wherein the customized input manner comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame; the object description information is used to determine a target of interest of the user; An image processing end configured to collect image frames; A cloud platform configured to receive the object description information sent by the client and receive image frames sent by the image processing end; and detect whether the target of interest exists in the image frames based on the object description information, and send a detection result to the image processing end; The image processing end is further configured to control a direction of image collection based on the detection result sent by the cloud platform, so that the target of interest is located at a preset position in the image frame; The image processing end is specifically configured to: if the image frame obtained at the first detection time instant contains the target of interest, calculate a moving speed of the target of interest in the direction of the imaging plane at the first detection time instant as a first moving speed based on a pixel position of the target of interest in the image frame; calculate a moving speed of the target of interest in the direction of the imaging plane at a third detection time instant as a third moving speed according to a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time instant before the first detection time instant and the first moving speed, wherein the third detection time instant is a next detection time instant of the second detection time instant, the second detection time instant is a next detection time instant of the first detection time instant, and a time interval between every two adjacent detection time instants is preset according to an execution time required for performing one speed prediction process; when the second detection time instant is reached, determine a peak speed required by a motor in the image acquisition device to reach at a specified time instant between the second detection time instant and the third detection time instant according to the third moving speed, a current speed of the motor at the current time instant, and a maximum speed supported by the motor; take the current speed of the motor at the current time instant as an initial speed, and adjust the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the peak speed at the specified time instant; when the specified time instant is reached, if the current speed of the motor at the current time instant does not reach the maximum speed supported, take the current speed of the motor at the current time instant as an initial speed, and adjust the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the third moving speed at the third detection time instant; the image processing end is further configured to: when the specified time instant is reached, if the current speed of the motor at the current time instant is the maximum speed supported, control the motor to keep the current speed until the third detection time instant.

14. An open self-defined detection target intelligent image processing device, characterized in that, The device comprises: a description information acquisition module configured to acquire object description information input by a user in a self-defined manner, wherein the self-defined input manner comprises at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interactive instruction containing position information in an image frame, and the object description information is used to determine a target of interest of the user; a detection module configured to detect whether the target of interest exists in an image frame collected based on the object description information; a first control module configured to control a direction of image acquisition so that the target of interest is located at a preset position in the image frame if it is detected that the target of interest exists in the image frame; the first control module comprises: a first speed calculation sub-module configured to calculate a moving speed of the target of interest in the direction of the imaging plane at the first detection time instant as a first moving speed based on a pixel position of the target of interest in the image frame if the image frame obtained at the first detection time instant contains the target of interest. The second speed calculation submodule is configured to calculate a third movement speed of the target of interest in the direction of the imaging plane at a third detection time according to a second movement speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time and the first movement speed, wherein the third detection time is a next detection time of the second detection time, the second detection time is a next detection time of the first detection time, and a time interval between every two adjacent detection times is preset according to an execution time required for performing one speed prediction process; The control submodule comprises: The peak speed determination unit is configured to determine a peak speed required by the motor at a specified time between the second detection time and the third detection time according to the third movement speed, a current speed of the motor at the current time, and a maximum speed supported by the motor when the second detection time is reached. The first speed adjustment unit is configured to adjust the speed of the motor in an increasing or decreasing manner with the current speed of the motor at the current time as an initial speed, so that the speed of the motor reaches the peak speed at the specified time. The second speed adjustment unit is configured to adjust the speed of the motor in an increasing or decreasing manner with the current speed of the motor at the current time as an initial speed when the specified time is reached, so that the speed of the motor reaches the third movement speed at the third detection time, if the current speed of the motor at the current time does not reach the maximum speed supported. The device further comprises: The second control module is configured to control the motor to maintain the current speed until the third detection time if the current speed of the motor at the current time is the maximum speed supported when the specified time is reached.

15. The apparatus of claim 14, wherein, The data format of the object description information is one of text, image, and audio. The detection module comprises: The first detection submodule is configured to, for one image frame collected, if the target of interest is detected not to exist in a continuous specified number of image frames before the image frame, or the image frame is the first image frame collected after the object description information is received, detect whether the target of interest exists in the image frame based on the object description information and the image frame by using a pre-trained open-vocabulary target detection model. The second detection submodule is configured to, if at least one of the continuous specified number of image frames is detected to have the target of interest, detect the image frame by using the open-vocabulary target detection model, and determine whether the target of interest exists in the image frame according to a pixel position of the detected target and pixel positions of the target of interest detected in the continuous specified number of image frames. And / or, The second detection submodule is specifically configured to: input the object description information and the image frame into a pre-trained open-vocabulary object detection model to obtain a pixel position of an object in the image frame that matches the object description information; or input the image frame and a preset prompt word into the pre-trained open-vocabulary object detection model to obtain a pixel position of an object in the image frame that matches the preset prompt word; and / or, the object description information is triggered by a user in a specified image frame; the detection module comprises: a third detection submodule configured to obtain a pixel position of each object in the specified image frame by using a pre-trained open-vocabulary object detection model, and determine whether the target of interest exists in the image frame according to the determined pixel position of each object and position information contained in the object description information; a fourth detection submodule configured to, for each image frame after the specified image frame, if it is detected that the target of interest exists in at least one image frame among a continuous specified number of image frames before the image frame, obtain a pixel position of each object in the image frame by using the open-vocabulary object detection model, and determine whether the target of interest exists in the image frame according to the determined pixel position of each object and the pixel position of the target of interest detected in the continuous specified number of image frames; and / or, the second speed calculation submodule is specifically configured to: determine a pixel distance between the target of interest and a reference stationary object in the image frame; input the determined pixel distance, a second moving speed of the target of interest in the direction of the imaging plane at a historical detection time before the first detection time, and the first moving speed into a pre-trained speed prediction model to obtain a moving speed of the target of interest in the direction of the imaging plane at the third detection time as a third moving speed; wherein the speed prediction model is trained based on sample data and corresponding sample labels; the sample data comprises a pixel distance between a sample object and a reference stationary object in a sample image frame collected at a detection time, a moving speed of the sample object in the direction of the imaging plane at a historical detection time before the detection time, and a moving speed of the sample object in the direction of the imaging plane at the detection time; and a sample label corresponding to one sample data comprises a speed of the sample object at a second detection time after the detection time; and / or, the first speed calculation submodule is specifically configured to: in a case where an image quality of an image region occupied by the target of interest in the image frame meets a preset condition, calculate a moving speed of the target of interest in the direction of the imaging plane at the first detection time as a first moving speed based on the pixel position of the target of interest in the image frame; wherein the preset condition comprises at least one of the following: the image region is not overexposed, the image region is not too dark, and the image region is clear; and / or, the peak rotational speed determination unit is specifically configured to: When the second detection time is reached, a preset formula is used to determine a peak speed of the motor required to be reached at a specified time between the second detection time and the third detection time according to the third moving speed, a current speed of the motor at the specified time, and a maximum speed supported by the motor. The preset formula is expressed as: ; represents the peak rotational speed; represents the maximum rotational speed supported by the motor; represents the third movement speed; represents the rotational speed of the motor at the second detection time.

16. An electronic device, comprising: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as:

17. A computer program product, characterised in that, The preset formula is expressed as:

18. A computer-readable storage medium, characterized in that, The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The preset formula is expressed as: The

Citation Information

Patent Citations

  • Video monitoring system and method for target detection and tracking

    CN108055501A

  • Object recognition

    US20220392209A1