Intelligent image processing method, device and equipment for open type custom detection target
By obtaining user-defined object description information and object detection model, controlling the direction of the image acquisition device and the rotation speed of the lens, the problem that existing equipment cannot capture multiple types of targets according to user needs is solved, and efficient personalized image acquisition is achieved.
Patent Information
- Application Number
- CN202510897125.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing image acquisition devices cannot continue to shoot non-preset targets according to users’ personalized needs, such as being unable to continuously shoot vehicles, pedestrians and animals at the same time, resulting in low availability and intelligence.
By obtaining user-defined object description information, a pre-trained open vocabulary object detection model is used to detect whether there is an interest target in the image frame, and the direction of the image acquisition device is controlled, so that the interest target is located at a preset position, and the speed prediction model is combined to adjust the lens rotation speed to follow the target, so as to achieve continuous shooting.
The image acquisition equipment continuously shoots the targets of interest based on user needs, improves the usability and intelligence of the equipment, and supports personalized shooting needs of multiple types of targets.
Smart Images

Figure CN120416668A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to an intelligent image processing method, device, and equipment for open custom detection targets. Background Art
[0002] Generally, the field of view of a camera is fixed. When using such a camera to perform image acquisition operations on a moving object, once the moving object moves out of the established field of view of the camera, the camera cannot capture an image of the moving object. To achieve continuous and effective image acquisition of a moving object, the camera can be mounted on a pan-tilt head. In this way, when controlling the rotation of the pan-tilt head, the pan-tilt head can drive the camera to rotate.
[0003] However, the types of targets that the above image acquisition device supports for continuous shooting are usually preset at the factory and cannot be continuously shot according to the personalized needs of users. For example, if a target supported by an image acquisition device for continuous shooting is set to a vehicle at the factory, then the image acquisition device only supports continuous shooting of vehicles and cannot achieve continuous shooting of other targets (pedestrians and animals) other than vehicles. This results in low usability and intelligence of the image acquisition device. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide an intelligent image processing method, device, and equipment for open custom detection targets to improve the usability and intelligence of image acquisition devices. The specific technical solutions are as follows:
[0005] In the first aspect of the embodiments of the present application, first, an intelligent image processing method for open custom detection targets is provided. The method includes:
[0006] Obtain object description information input by the user customarily; wherein, the custom input methods include at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in an image frame; the object description information is used to determine the user's interested target;
[0007] Based on the object description information, detect whether the interested target exists in the captured image frame;
[0008] If the interested target is detected in the image frame, control the direction of image acquisition so that the interested target is located at a preset position in the image frame.
[0009] In some embodiments, the data format of the object description information is one of text, image, and audio;
[0010] Detecting whether the target of interest exists in the acquired image frame based on the object description information includes:
[0011] For an acquired image frame, if it is detected that the target of interest does not exist in any of the consecutive specified number of image frames preceding this image frame, or if this image frame is the first image frame acquired after receiving the object description information, then based on the object description information and this image frame, use a pre-trained open-vocabulary object detection model to detect whether the target of interest exists in this image frame;
[0012] If it is detected that at least one of the consecutive specified number of image frames has the target of interest, then use the open-vocabulary object detection model to detect this image frame; and determine whether the target of interest exists in this image frame according to the pixel positions of the detected targets and the pixel positions of the targets of interest detected in the consecutive specified number of image frames.
[0013] In some embodiments, using the open-vocabulary object detection model to detect this image frame includes:
[0014] Input the object description information and this image frame into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets in this image frame that conform to the object description information;
[0015] Or,
[0016] Input this image frame and a preset prompt word into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets in this image frame that conform to the preset prompt word.
[0017] In some embodiments, the object description information is triggered by the user in a specified image frame;
[0018] Detecting whether the target of interest exists in the acquired image frame based on the object description information includes:
[0019] Use a pre-trained open-vocabulary object detection model to obtain the pixel positions of each target in the specified image frame, and determine whether the target of interest exists in this image frame according to the determined pixel positions of each target and the position information included in the object description information;
[0020] For each image frame after the specified image frame, if it is detected that at least one of the consecutive specified number of image frames before this image frame has the target of interest, use the open-vocabulary object detection model to obtain the pixel positions of each target in this image frame, and based on the determined pixel positions of each target, and the pixel positions of the targets of interest detected in the consecutive specified number of image frames, determine whether the target of interest exists in this image frame.
[0021] In some embodiments, the step of, if it is detected that the target of interest exists in the image frame, controlling the direction of image acquisition so that the target of interest is located at a preset position in the image frame, includes:
[0022] If the target of interest exists in the image frame obtained at the first detection moment, based on the pixel position of the target of interest in this image frame, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed;
[0023] According to the second moving speed and the first moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment, calculate the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed; wherein, the third detection moment is the next detection moment of the second detection moment; the second detection moment is the next detection moment of the first detection moment;
[0024] When the second detection moment is reached, adjust the rotation speed of the motor according to the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens.
[0025] In some embodiments, the step of calculating the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed according to the second moving speed and the first moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment includes:
[0026] Determine the pixel distance between the target of interest and a reference stationary object in this image frame;
[0027] Input the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment, and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment, which is used as the third moving speed;
[0028] Wherein, the speed prediction model is: trained based on sample data and corresponding sample labels;
[0029] The sample data includes: the pixel distance between the sample object and the reference stationary object in the sample image frame collected at a detection moment, the moving speed of the sample object in the direction of the imaging plane at a historical detection moment before this detection moment, and the moving speed of the sample object in the direction of the imaging plane at this detection moment; the sample label corresponding to a sample data includes: the speed of the sample object at the second detection moment after this detection moment.
[0030] In some embodiments, calculating the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed based on the pixel position of the target of interest in the image frame includes:
[0031] When the image quality of the image area occupied by the target of interest in the image frame meets the preset conditions, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment based on the pixel position of the target of interest in the image frame, and use it as the first moving speed;
[0032] Wherein, the preset conditions include at least one of the following: the image area is not overexposed, the image area is not underexposed, and the image area is clear.
[0033] In some embodiments, when reaching the second detection moment, adjust the rotation speed of the motor according to the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens, including:
[0034] When reaching the second detection moment, determine the peak rotation speed that the motor needs to reach at a specified moment between the second detection moment and the third detection moment according to the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor;
[0035] Taking the rotation speed of the motor at the current moment as the initial rotation speed, adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the peak rotation speed at the specified moment;
[0036] When the specified moment is reached, if the rotation speed of the motor at the current moment does not reach the maximum rotation speed supported, taking the rotation speed of the motor at the current moment as the initial rotation speed, adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the third moving speed at the third detection moment;
[0037] The method further includes:
[0038] When the specified moment is reached, if the rotation speed of the motor at the current moment is the maximum rotation speed supported, control the motor to maintain the rotation speed at the current moment until the third detection moment.
[0039] In some embodiments, when the second detection moment is reached, according to the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor, determining the peak rotation speed required by the motor at a specified moment between the second detection moment and the third detection moment includes:
[0040] When the second detection moment is reached, according to a preset formula, based on the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor, determine the peak rotation speed required by the motor at a specified moment between the second detection moment and the third detection moment;
[0041] Among them, the preset formula is expressed as:
[0042]
[0043] represents the peak rotation speed; is the maximum rotation speed supported by the motor; represents the third moving speed; represents the rotation speed of the motor at the second detection moment.
[0044] In the second aspect of the embodiments of the present application, there is provided an intelligent image processing device for an open custom detection target, the device includes:
[0045] A lens for collecting image frames;
[0046] A processor for executing the intelligent image processing method for the open custom detection target described in any one of the above.
[0047] In some embodiments, the device further includes: a pan-tilt head;
[0048] The processor is specifically configured to send a control instruction to the pan-tilt head;
[0049] The pan-tilt head is configured to adjust the image acquisition direction of the lens according to the control instruction sent by the processor, so that the target of interest is located at a preset position in the image frame.
[0050] In a third aspect of the embodiments of the present application, there is provided an intelligent image processing system for open-ended custom detection of targets, the system includes:
[0051] A client, configured to receive object description information custom-input by a user; wherein, the custom input method includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in an image frame; the object description information is used to determine the target of interest of the user;
[0052] An image processing terminal, configured to collect an image frame, and based on the object description information, detect whether the target of interest exists in the collected image frame; if the target of interest is detected in the image frame, control the image acquisition direction so that the target of interest is located at a preset position in the image frame.
[0053] In some embodiments, the client is integrated in the image processing terminal.
[0054] In some embodiments, the system further includes: a cloud platform;
[0055] The cloud platform is configured to receive the object description information sent by the client and forward it to the image processing terminal.
[0056] In a fourth aspect of the embodiments of the present application, there is provided an intelligent image processing system for open-ended custom detection of targets, the system includes:
[0057] A client, configured to receive object description information custom-input by a user; wherein, the custom input method includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in an image frame; the object description information is used to determine the target of interest of the user;
[0058] An image processing terminal, configured to collect an image frame;
[0059] A cloud platform, configured to receive the object description information sent by the client and the image frame sent by the image processing terminal; and based on the object description information, detect whether the target of interest exists in the image frame and send a detection result to the image processing terminal;
[0060] The image processing terminal is further configured to control the direction of image acquisition according to the detection result sent by the cloud platform, so that the target of interest is located at a preset position in the image frame.
[0061] In a fifth aspect of the embodiments of the present application, an intelligent image processing device for open custom detection targets is provided. The device includes:
[0062] A description information acquisition module, configured to acquire object description information custom-input by a user; wherein, the custom input method includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in the image frame; the object description information is used to determine the target of interest of the user;
[0063] A detection module, configured to detect whether the target of interest exists in the acquired image frame based on the object description information;
[0064] A first control module, configured to control the direction of image acquisition if the target of interest is detected in the image frame, so that the target of interest is located at a preset position in the image frame.
[0065] In some embodiments, the data format of the object description information is one of text, image, and audio;
[0066] The detection module includes:
[0067] A first detection sub-module, configured to, for an acquired image frame, if it is detected that the target of interest does not exist in any of the consecutive specified number of image frames preceding the current image frame, or if the current image frame is the first image frame acquired after receiving the object description information, detect whether the target of interest exists in the current image frame based on the object description information and the current image frame by using a pre-trained open vocabulary object detection model;
[0068] A second detection sub-module, configured to, if it is detected that at least one of the consecutive specified number of image frames has the target of interest, use the open vocabulary object detection model to detect the current image frame; and determine whether the target of interest exists in the current image frame according to the pixel positions of the detected targets and the pixel positions of the targets of interest detected in the consecutive specified number of image frames.
[0069] In some embodiments, the second detection sub-module is specifically configured to:
[0070] Input the object description information and the current image frame into a pre-trained open vocabulary object detection model to obtain the pixel positions of the targets in the current image frame that conform to the object description information;
[0071] Alternatively,
[0072] input the image frame and the preset prompt words into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the objects in the image frame that match the preset prompt words.
[0073] In some embodiments, the object description information is triggered by the user in a specified image frame;
[0074] The detection module includes:
[0075] A third detection sub-module, configured to use a pre-trained open-vocabulary object detection model to obtain the pixel positions of each object in the specified image frame, and determine whether the target of interest exists in the image frame according to the determined pixel positions of each object and the position information included in the object description information;
[0076] A fourth detection sub-module, for each image frame after the specified image frame, if it is detected that at least one image frame among a continuous specified number of image frames before this image frame has the target of interest, use the open-vocabulary object detection model to obtain the pixel positions of each object in this image frame, and determine whether the target of interest exists in this image frame according to the determined pixel positions of each object and the pixel positions of the target of interest detected in the continuous specified number of image frames.
[0077] In some embodiments, the first control module includes:
[0078] A first speed calculation sub-module, configured to, if the target of interest exists in the image frame obtained at the first detection moment, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment based on the pixel position of the target of interest in this image frame, as the first moving speed;
[0079] A second speed calculation sub-module, configured to calculate the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed according to the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment and the first moving speed; wherein, the third detection moment is the next detection moment of the second detection moment; the second detection moment is the next detection moment of the first detection moment;
[0080] A control sub-module, configured to, when the second detection moment is reached, adjust the rotation speed of the motor according to the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens.
[0081] In some embodiments, the second speed calculation sub-module is specifically configured to:
[0082] Determine the pixel distance between the target of interest and the reference static object in the image frame;
[0083] Input the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment, and the first moving speed into a pre-trained speed prediction model, and obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed;
[0084] Wherein, the speed prediction model is trained based on sample data and corresponding sample labels;
[0085] The sample data includes: the pixel distance between the sample object and the reference static object in the sample image frame collected at a detection moment, the moving speed of the sample object in the direction of the imaging plane at the historical detection moment before this detection moment, and the moving speed of the sample object in the direction of the imaging plane at this detection moment; a sample label corresponding to a sample data includes: the speed of the sample object at the second detection moment after this detection moment.
[0086] In some embodiments, the first speed calculation sub-module is specifically configured to:
[0087] When the image quality of the image region occupied by the target of interest in the image frame meets a preset condition, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment based on the pixel position of the target of interest in the image frame as the first moving speed;
[0088] Wherein, the preset condition includes at least one of the following: the image region is not overexposed, the image region is not underexposed, and the image region is clear.
[0089] In some embodiments, the control sub-module is specifically configured to:
[0090] A peak rotation speed determination unit, configured to, when reaching the second detection moment, determine the peak rotation speed required for the motor at a specified moment between the second detection moment and the third detection moment according to the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor;
[0091] A first rotation speed adjustment unit, configured to use the rotation speed of the motor at the current moment as the initial rotation speed, and adjust the rotation speed of the motor in an increasing or decreasing manner, so that the rotation speed of the motor reaches the peak rotation speed at the specified moment;
[0092] A second rotation speed adjustment unit, configured to, when reaching the specified moment, if the rotation speed of the motor at the current moment does not reach the maximum rotation speed supported, use the rotation speed of the motor at the current moment as the initial rotation speed, and adjust the rotation speed of the motor in an increasing or decreasing manner, so that the rotation speed of the motor reaches the third moving speed at the third detection moment;
[0093] The device further includes:
[0094] A second control module, configured to, when reaching the specified moment, if the rotation speed of the motor at the current moment is the maximum rotation speed supported, control the motor to maintain the rotation speed at the current moment until the third detection moment.
[0095] In some embodiments, the peak rotation speed determination unit is specifically configured to:
[0096] When reaching the second detection moment, according to a preset formula, determine the peak rotation speed required for the motor at a specified moment between the second detection moment and the third detection moment according to the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor;
[0097] Wherein, the preset formula is expressed as:
[0098]
[0099] represents the peak rotation speed; is the maximum rotation speed supported by the motor; represents the third moving speed; represents the rotation speed of the motor at the second detection moment.
[0100] Another aspect of the embodiments of the present application provides an electronic device, including:
[0101] A memory, configured to store a computer program;
[0102] A processor, when executing a program stored in a memory, implements the intelligent image processing method for open-ended custom detection targets described above.
[0103] In another aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the intelligent image processing method for open-ended custom detection targets described above is implemented.
[0104] In another aspect of the embodiments of the present application, a computer program product containing instructions is provided. When it runs on a computer, it causes the computer to execute the intelligent image processing method for open-ended custom detection targets described above.
[0105] Advantages of the embodiments of the present application:
[0106] Based on the intelligent image processing method for open-ended custom detection targets provided by the embodiments of the present application, it supports users to input object description information for determining the target of interest in a custom input manner. Correspondingly, based on the object description information, it can be detected whether there is a target of interest in the captured image frame. Furthermore, if a target of interest is detected in the image frame, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frame. That is to say, the electronic device can control the image acquisition device to continuously capture the target (i.e., the target of interest) that conforms to the object description information in the image frame according to the actual needs of the user, improving the usability and intelligence level of the image acquisition device.
[0107] Of course, implementing any product or method of the present application does not necessarily require achieving all the above-mentioned advantages simultaneously. Description of the Drawings
[0108] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0109] Figure 1 It is the first flowchart of the intelligent image processing method for open-ended custom detection targets provided by the embodiments of the present application;
[0110] Figure 2 It is the first flowchart for control after detecting a target of interest provided by the embodiments of the present application;
[0111] Figure 3 It is the second flowchart for control after detecting a target of interest provided by the embodiments of the present application;
[0112] Figure 4 This is the third flowchart for control after detecting an object of interest provided by an embodiment of the present application;
[0113] Figure 5 This is a schematic flowchart of the control of an image acquisition device provided by an embodiment of the present application;
[0114] Figure 6 This is a schematic structural diagram of an intelligent image processing device provided by an embodiment of the present application;
[0115] Figure 7 This is a schematic structural diagram of an intelligent image processing system provided by an embodiment of the present application;
[0116] Figure 8 This is a schematic structural diagram of another intelligent image processing system provided by an embodiment of the present application;
[0117] Figure 9 This is a structural diagram of an intelligent image processing device with an open custom detection target provided by an embodiment of the present application;
[0118] Figure 10 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0119] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0120] To solve the technical problems of low availability and low intelligence of image acquisition devices, an embodiment of the present application provides an intelligent image processing method with an open custom detection target, and this method can be applied to an electronic device.
[0121] Among them, the description of the intelligent image processing device with an open custom detection target can be seen in the subsequent embodiments.
[0122] See Figure 1 , Figure 1 This is the first flowchart of the intelligent image processing method with an open custom detection target provided by an embodiment of the present application, and this method includes the following steps:
[0123] S101: Obtain the object description information input by the user customarily.
[0124] Among them, the custom input methods include at least one of the following: inputting text, inputting images, inputting audio, and triggering an interaction instruction containing position information in the image frame; the object description information is used to determine the user's target of interest.
[0125] S102: Based on the object description information, detect whether there is a target of interest in the captured image frame.
[0126] S103: If a target of interest is detected in the image frame, control the direction of image acquisition so that the target of interest is located at a preset position in the image frame.
[0127] Based on the intelligent image processing method for open custom detection targets provided in the embodiments of the present application, it supports the user to input object description information for determining the user's target of interest in a custom input manner. Correspondingly, based on the object description information, it can be detected whether there is a target of interest in the captured image frame. Furthermore, if a target of interest is detected in the image frame, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frame. That is to say, the electronic device can control the image acquisition device to continuously capture the target (i.e., the target of interest) in the image frame that conforms to the object description information according to the actual needs of the user, improving the usability of the image acquisition device.
[0128] Regarding step S101, the user can input object description information for determining the target of interest in a custom input manner.
[0129] Among them, the custom input methods include at least one of the following: inputting text, inputting images, inputting audio, and triggering an interaction instruction containing position information in the image frame.
[0130] In one implementation, the electronic device can provide an interaction interface for receiving the object description information input by the user. Alternatively, a client independent of the electronic device can provide an interaction interface for receiving the object description information input by the user, and the client can communicate with the electronic device. Correspondingly, the user can input the object description information through the interaction interface of the client, and then the client can send the object description information input by the user to the electronic device.
[0131] The custom input methods include:
[0132] Method 1: Inputting text.
[0133] In this method, the user can input text representing the target to be continuously captured through the text box in the interaction interface, and click the submit button. Then, the electronic device can obtain the text input by the user as the object description information. That is, the data format of the object description information is text.
[0134] For example, if the target for continuous shooting is a person wearing a white coat, the user can input the text "person wearing a white coat".
[0135] Method 2: Input an image.
[0136] The user can upload, through the image upload component in the interaction interface, an image containing the target for continuous shooting. Furthermore, the electronic device can obtain the image input by the user as object description information. That is, the data format of the object description information is an image.
[0137] For example, if the target for continuous shooting is a tabby cat, the user can input an image containing the tabby cat.
[0138] Method 3: Input an audio.
[0139] The user can upload, through the audio upload component in the interaction interface, an audio representing the target for continuous shooting. Furthermore, the electronic device can obtain the audio input by the user as object description information. At this time, the data format of the object description information is an audio.
[0140] For example, if the target for continuous shooting is a vehicle of a certain brand, the user can upload an audio with the content "vehicle of a certain brand".
[0141] Method 4: Trigger an interaction instruction containing position information in an image frame.
[0142] The electronic device can display the image frames captured in real time by the image capture device in the interaction interface. If there is a target in the currently displayed image frame that the user needs to continuously shoot, the user can, in this image frame, trigger an interaction instruction containing position information at the position of the target to be continuously shot in this image frame by means of clicking or sliding selection. Correspondingly, the electronic device can obtain the interaction instruction containing position information triggered by the user in the image frame.
[0143] It can be understood that since the target to be continuously shot may be in a moving state, correspondingly, the position of the same target in the image frames captured at different times may also be different. Therefore, the object description information input by the user by triggering the interaction instruction only applies to the image frame (which can be called the specified image frame of this interaction instruction) corresponding to the triggered interaction instruction. That is, the electronic device can only detect whether there is an object of interest in the specified image frame of this interaction instruction based on this interaction instruction. It cannot detect whether there is an object of interest in other image frames except the specified image frame of this interaction instruction based on this interaction instruction.
[0144] Regarding step S102, starting from the moment when the electronic device obtains the object description information input by the user, the electronic device can detect whether there is an object of interest in the captured image frames.
[0145] Among them, for any captured image frame, if the detection result of the image frame indicates that there is an object of interest in the image frame, the electronic device can also determine the pixel position of the object of interest in the image frame.
[0146] When an object (i.e., the object of interest) that conforms to the object description information enters the field of view of the image acquisition device, the electronic device can start continuously shooting the object of interest. That is, the electronic device can control the direction of image acquisition so that the object of interest is located at a preset position in the image frame. Among them, the specific process of controlling the direction of image acquisition will be described in subsequent embodiments.
[0147] It can be understood that since the range that the direction of image acquisition can adjust is limited, the maximum field of view that the image acquisition device can capture is also limited. For an object of interest in a moving state, as the object of interest moves continuously, the object of interest may leave the maximum field of view that the image acquisition device can capture. Correspondingly, the image acquisition device can stop continuously shooting the object of interest.
[0148] In one implementation, after the image acquisition device stops continuously shooting the object of interest, the electronic device can adjust the direction of the image acquisition device for capturing images to the initialization direction. Among them, the electronic device can control the rotation of the motor in the image acquisition device to adjust the direction of the image acquisition device for capturing images. The specific process of adjusting the direction of the image acquisition device for capturing images will be described in subsequent embodiments.
[0149] If the data format of the object description information is one of text, image, and audio, then after an object of interest leaves the maximum field of view that the image acquisition device can capture for a period of time, the image acquisition device can re-detect whether there is an object (new object of interest) that conforms to the object description information in the captured image frame, and after detecting the new object of interest, continuously shoot the new object of interest.
[0150] In one implementation, after each new object of interest is determined, the electronic device can assign a unique identifier (ID, Identity document) corresponding to the object of interest to the object of interest. That is, the identifiers corresponding to different objects of interest are also different. For example, the unique identifier corresponding to an object of interest can represent the time sequence in which the object of interest appears in the image acquisition device. Subsequently, the user can view the video segments corresponding to each object of interest through the unique identifiers corresponding to each object of interest.
[0151] In some embodiments, the data format of the object description information is one of text, image, and audio.
[0152] Correspondingly, step S102 includes:
[0153] Step S1021: For a captured image frame, if it is detected that there is no target of interest in any of the consecutive specified number of image frames preceding this image frame, or if this image frame is the first image frame captured after receiving the object description information, then based on the object description information and this image frame, use a pre-trained open-vocabulary object detection model to detect whether there is a target of interest in this image frame.
[0154] Step S1022: If it is detected that there is at least one image frame among the consecutive specified number of image frames that has a target of interest, then use the open-vocabulary object detection model to detect this image frame; and based on the pixel positions of the detected targets and the pixel positions of the targets of interest detected in the consecutive specified number of image frames, determine whether there is a target of interest in this image frame.
[0155] In the embodiment of the present application, for step S1021, for a captured image frame, if this image frame is the first image frame captured after receiving the object description information, it indicates that: the electronic device needs to start from this image frame to detect whether there is a target of interest within the field of view of the image acquisition device. Correspondingly, based on the obtained object description information and this first image frame, a pre-trained open-vocabulary object detection model can be used to detect whether there is a target of interest in this image frame.
[0156] Or, for a captured image frame, if it is detected that there is no target of interest in any of the consecutive specified number of image frames preceding this image frame, it indicates that: within the historical time period corresponding to the consecutive specified number of image frames preceding this image frame, no target of interest has been detected within the field of view of the image acquisition device. This situation may be that no target has been detected since the beginning (i.e., since the moment of receiving the object description information), or it may be that there were image frames with detected targets of interest in the past, but as the target of interest moves, the target of interest has left the field of view of the image acquisition device. Since no target of interest has been detected within the field of view of the lens within the historical time period corresponding to the consecutive specified number of image frames preceding this image frame, at this time, the electronic device can re-determine the target of interest. That is, the electronic device can only based on the obtained object description information and this image frame, use a pre-trained open-vocabulary object detection model to detect whether there is a target of interest in this image frame.
[0157] Among them, the consecutive specified number can be determined by those skilled in the art according to actual needs. For example, the consecutive specified number can be 10.
[0158] Correspondingly, the electronic device can use the obtained object description information as a prompt (Prompt) to be input into the open-vocabulary object detection model. That is, the image frame and the prompt (i.e., the object description information) are input into the pre-trained open-vocabulary object detection model to obtain a detection result indicating whether there is an object of interest in the image frame.
[0159] Among them, the pre-trained open-vocabulary object detection model (OVOD, Open Vocabulary Object Detection) is an object detection model combined with a large language model (LLM, Large Language Model). For example, the open-vocabulary object detection model can be: YOLO-World (You Only Look Once - World, a real-time open-vocabulary object detection model), or DINO (Detection Transformer, an object detection model implemented based on a transformer network), etc.
[0160] Regarding step S1022, for a captured image frame, if it is detected that at least one of the consecutive specified number of image frames before the current image frame has an object of interest, it indicates that within the historical time period corresponding to the specified number of image frames before the current image frame, an object of interest has been detected within the field of view of the camera.
[0161] Correspondingly, in order to achieve continuous shooting of the same object of interest as much as possible, the electronic device can use the open-vocabulary object detection model to detect the image frame to obtain the pixel positions of the detected objects (which can be called candidate objects). Furthermore, by combining the pixel positions of the objects of interest detected in the consecutive specified number of image frames before the current image frame, the candidate objects that are objects of interest in the current image frame are determined.
[0162] In some embodiments, the pixel positions of the candidate objects in the current image frame can be obtained according to any one of the following methods a or b:
[0163] Method a: Input the object description information and the current image frame into the pre-trained open-vocabulary object detection model to obtain the pixel positions of the objects in the current image frame that match the object description information.
[0164] In the embodiments of the present application, the electronic device can input the object description information and the current image frame into the pre-trained open-vocabulary object detection model to obtain the pixel positions of all the objects in the current image frame that match the object description information (which can be called the first candidate objects). Among them, the specific process of using the pre-trained open-vocabulary object detection model for detection can refer to the above embodiments and will not be elaborated here.
[0165] Based on the above processing, it can be ensured that the detected targets (the first alternative targets) all conform to the object description information input by the user. Subsequently, determining the target of interest from the first alternative targets can improve the accuracy of the determined target of interest.
[0166] Method b: Input the image frame and the preset prompt word into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets in the image frame that conform to the preset prompt word.
[0167] In the embodiment of the present application, the electronic device can input the image frame and the preset prompt word into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets (which can be called the second alternative targets) in the image frame that conform to the preset prompt word.
[0168] Among them, the preset prompt word can be: a prompt word pre-set by a technician. For example, the preset prompt word can include at least one of the following: people, animals, and vehicles, etc.
[0169] Based on the above processing, the preset prompt word can be a prompt word with relatively high generality pre-set by a technician, which can ensure that various movable targets in the image frame can be detected as much as possible and avoid missing the detection of targets.
[0170] It can be understood that in the multiple consecutive image frames (the image frame and the consecutive specified number of image frames before the image frame), for the same target, the positions of the target in two adjacent image frames are usually close, and the position change of the target in the multiple consecutive image frames is usually continuous.
[0171] Therefore, after obtaining the pixel positions of the alternative targets detected in the image frame, the electronic device can determine whether there is a target of interest in the image frame by comparing the pixel positions of each alternative target with the pixel positions of the targets of interest detected in the consecutive specified number of image frames (which can be called the historical pixel positions of the targets of interest).
[0172] In one implementation, the electronic device can use a preset object tracking (OT, Object Tracking) algorithm to determine the pixel positions of the targets of interest in the image frame according to the pixel positions of each alternative target and the pixel positions of the targets of interest in the collected historical image frames (that is, the pixel positions of the targets of interest detected in the consecutive specified number of image frames). For example, the object tracking algorithm can be: the Hungarian matching algorithm.
[0173] Correspondingly, the electronic device can obtain the pixel positions of the target of interest in each historical image frame. For the pixel positions of each alternative target in this image frame, calculate the intersection over union (IOU) between the image area occupied by this alternative target in this image frame and the image areas occupied by the target of interest in each historical image frame, to obtain the cost matrix corresponding to this alternative target in this image frame. Among them, the larger the IOU corresponding to this alternative target in this image frame, the smaller the value in the corresponding cost matrix. Furthermore, among the alternative targets in this image frame, determine the alternative target with the smallest total cost represented by the corresponding cost matrix and not greater than the preset threshold as the target of interest in this image frame, and determine the pixel position of the target of interest in this image frame. If there is no alternative target whose corresponding cost matrix is not greater than the preset threshold, it can be determined that none of the alternative targets in this image frame are the targets of interest in the historical image frames, and furthermore, it can be determined that there is no target of interest in this image frame.
[0174] In another implementation, the electronic device can use a pre-trained target tracking model to determine the pixel position of the target of interest in this image frame according to the pixel positions of each alternative target included in this image frame and the pixel positions of the target of interest in the collected historical image frames.
[0175] For example, the target tracking model can be a deep learning model trained based on the DeepSORT (Deep Simple Online and Realtime Tracking) algorithm.
[0176] Correspondingly, the electronic device can input this image frame, the pixel positions of each alternative target in this image frame, each historical image frame, and the pixel positions of the target of interest in each historical image frame into the pre-trained target tracking model.
[0177] Correspondingly, the multi-target tracking model can extract: the image features of the image areas occupied by each alternative target in this image frame, and the image features of the image areas occupied by the target of interest in each historical image frame. And for each alternative target in this image frame, the intersection over union between the image area occupied by this alternative target and the image areas occupied by the target of interest in each historical image frame. Furthermore, combine the similarity between the image features of the image areas occupied by each alternative target and the image features of the image areas occupied by the target of interest in each historical image frame, and the calculated intersection over union for matching to determine whether there is a target of interest among the alternative targets in this image frame.
[0178] In this way, the target tracking model uses the image features and pixel positions of each alternative target in the image frame, and performs multimodal matching with the image features and pixel positions of the target of interest in each historical image frame. In this way, the robustness of the matching can be improved, especially suitable for the case where the alternative targets in the image frame are occluded. When it is detected that at least one image frame among a continuously specified number of image frames has a target of interest, the electronic device can determine whether there is a target of interest in the image frame to be processed according to the pixel positions of each alternative target in the image frame and the pixel positions of the target of interest in the collected historical image frames. In this way, the positioning of the target of interest in the image frame to be processed is realized, ensuring the continuity of the detection of the same target.
[0179] Based on the above processing, when the data format of the object description information is one of text, image, and audio, the electronic device can determine the targets of interest that appear in different time periods. That is, the electronic device can detect whether there is a new target of interest in the collected image frames after a period of time when a target of interest leaves the field of view of the image acquisition device. In addition, the user can customize the description information of the target of interest input according to actual needs, and use the open vocabulary object detection model to detect the objects that meet the description information in the image frame as the targets of interest. In this way, the categories of objects that support continuous shooting can be expanded to meet the personalized needs of users in different scenarios.
[0180] If the object description information is triggered by the user in a specified image frame, it indicates that the object description information is only used to describe the target of interest in the specified image frame.
[0181] In some embodiments, the object description information is triggered by the user in a specified image frame. Correspondingly, step S102 includes:
[0182] Step S102a: Use the pre-trained open vocabulary object detection model to obtain the pixel positions of each target in the specified image frame, and determine whether there is a target of interest in the image frame according to the determined pixel positions of each target and the position information included in the object description information.
[0183] Step S102b: For each image frame after the specified image frame, if it is detected that at least one image frame among the continuously specified number of image frames before the image frame has a target of interest, use the open vocabulary object detection model to obtain the pixel positions of each target in the image frame, and determine whether there is a target of interest in the image frame according to the determined pixel positions of each target and the pixel positions of the targets of interest detected in the continuously specified number of image frames.
[0184] In an embodiment of the present application, when a user triggers an interaction instruction containing location information in an image frame (i.e., a specified image frame), correspondingly, the electronic device can utilize the location information contained in the interaction instruction to determine an object of interest in the specified image frame.
[0185] For step S102a, the electronic device can input the image frame and a preset prompt word into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the objects in the image frame that conform to the preset prompt word.
[0186] Among them, the preset prompt word can be: a prompt word pre-set by a technician. For example, the preset prompt word can include at least one of the following: people, animals, and vehicles, etc.
[0187] In this way, the preset prompt word can be a prompt word with relatively high generality pre-set by a technician, which can ensure that various movable objects in the image frame can be detected as much as possible, avoiding the omission of objects.
[0188] It can be understood that if the location information contained in the interaction instruction represents a pixel coordinate, the electronic device can determine, among the detected objects, the object whose image area contains this pixel coordinate as the object of interest.
[0189] If this pixel coordinate does not belong to the image area occupied by any of the detected objects, the object in the specified image frame that is closest to this pixel coordinate can be determined as the object of interest. Among them, the distance between an object and this pixel coordinate can be expressed as: the closest distance between each pixel coordinate in the image area occupied by this object and this pixel coordinate.
[0190] If the location information contained in the interaction instruction represents a piece of image area in the specified image frame, the electronic device can determine the object in the specified image frame with the highest intersection-over-union ratio with this image area as the object of interest. Among them, the method of calculating the intersection-over-union ratio can refer to the above embodiment and will not be elaborated here.
[0191] For step S102b, for each image frame after the specified image frame, the process of determining whether there is an object of interest in this image frame can refer to the process of performing detection in the above embodiment, that is, if at least one of the continuous specified number of image frames before this image frame is detected to have an object of interest, the process will not be elaborated here.
[0192] Based on the above processing, the user can determine the location information of the object of interest in the image to be processed by clicking or bounding box selection. Furthermore, the electronic device can determine the object of interest that needs to be continuously captured from this image frame through this location information. In this way, the operation steps of the user can be simplified and the user experience can be improved.
[0193] Regarding step S103, for a captured image frame, if an object of interest is detected in the image frame, the electronic device can control the image capture direction so that the object of interest is located at a preset position in the image frame.
[0194] For example, the preset position can be the position of the center point in the image frame.
[0195] It can be understood that usually the image capture device captures image frames at a high frequency, and the difference between two adjacent captured image frames is small. To reduce the computational complexity of continuously tracking the target, the electronic device can perform the detection processes of the above steps S101 to S102 on the image frames captured at each detection moment according to a preset detection frequency.
[0196] In this way, it is avoided to detect each captured image frame, reducing the demand pressure on the computing resources of the electronic device.
[0197] Correspondingly, for a detection moment, an image frame captured at this detection moment (which can be called the image frame to be processed). If there is an object of interest in the image frame to be processed, the detection moment corresponding to this image frame to be processed (i.e., the first detection moment) is the first detection moment during the process of continuously shooting the object of interest by the image capture device.
[0198] See Figure 2 , Figure 2 which is the first flowchart for control after detecting an object of interest provided by the embodiments of the present application.
[0199] S201: If there is an object of interest in the image frame obtained at the first detection moment, based on the pixel position of the object of interest in the image frame, calculate the moving speed of the object of interest in the direction of the imaging plane at the first detection moment as the first moving speed.
[0200] S202: According to the second moving speed and the first moving speed of the object of interest in the direction of the imaging plane at the historical detection moment before the first detection moment, calculate the moving speed of the object of interest in the direction of the imaging plane at the third detection moment as the third moving speed.
[0201] Among them, the third detection moment is the next detection moment of the second detection moment; the second detection moment is the next detection moment of the first detection moment.
[0202] S203: When the second detection moment is reached, adjust the rotational speed of the motor according to the third moving speed, the rotational speed of the motor at the current moment in the image acquisition device, and the maximum rotational speed supported by the motor, so that the rotational speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotational angle of the lens.
[0203] Based on the above processing, the first detection moment, the second detection moment, and the third detection moment are three consecutive adjacent detection moments. If there is a target of interest in the image frame obtained at the first detection moment (which can be called the image frame to be processed), then based on the pixel position of the target of interest in the image frame to be processed, the moving speed of the target of interest in the direction of the imaging plane at the first detection moment (i.e., the first moving speed) can be calculated. Similarly, for the historical detection moments before the first detection moment, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moments can also be calculated. Furthermore, the moving speed of the target of interest in the direction of the imaging plane at the third detection moment (i.e., the third moving speed) can be predicted according to the change trend of the moving speeds of the target of interest at different detection times. Correspondingly, when the second detection moment is reached, the rotational speed of the motor can be adjusted according to the obtained third moving speed, the rotational speed of the motor at the current moment in the image acquisition device, and the maximum rotational speed supported by the motor. Make the rotational speed of the lens in the image acquisition device at the third detection moment consistent with the third moving speed, that is to say, starting from the second detection moment, adjust the rotational speed of the lens so that at the third detection moment, the pixel position change speed of the target of interest in the image frame caused by the rotation of the lens in the image acquisition device is consistent with the pixel position change speed of the target of interest caused by its own movement, which can ensure the synchronization of the rotational speed of the lens in the image acquisition device and the moving speed of the target of interest, that is, the rotational speed of the lens can be adjusted according to the moving speed of the target of interest so that the lens in the image acquisition device can continuously shoot the moving target of interest. And, from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotational angle of the lens, that is to say, within the time interval from the second detection moment to the third detection moment, the pixel position of the target of interest moving in the picture caused by the rotation of the lens can be used to compensate the pixel position of the target of interest moving in the picture caused by its own movement, ensuring the synchronization of the rotational amount of the lens in the image acquisition device and the moving amount of the target of interest, that is, compensating the moving amount of the target of interest through the rotational amount of the lens so that the position of the moving target of interest in the image frame acquired within this time interval is stable. In this way, to a certain extent, the image acquisition device can continuously shoot the target of interest.
[0204] Regarding step S201, the first detection moment means: the detection moment when an object of interest exists in the image frame obtained at this detection moment.
[0205] After starting to continuously capture the object of interest through the image acquisition device, each time the electronic device reaches a detection moment, it executes the above steps S201 to S202 once to obtain the moving speed of the object of interest in the direction of the imaging plane at the second detection moment after this detection moment. Herein, the object of interest is the object that the image acquisition device needs to continuously capture, and the specific execution process of steps S201 to S202 will be described in subsequent embodiments. For the sake of convenience of description, the process of the electronic device executing the above steps S201 to S202 once can be called a speed prediction process.
[0206] The time interval between every two adjacent detection moments (which can be called the detection time interval) is preset by the technician according to the execution time required for executing a speed prediction process. For example, the set detection time interval is usually not less than the time consumption required for the electronic device to execute a speed prediction process according to steps S201 to S202. For instance, if the execution time required for executing a speed prediction process is about 120 milliseconds (i.e., 0.12 seconds), the detection time interval can be set to 120 milliseconds.
[0207] It can be understood that the detection moment may not be synchronized with the moment when the image acquisition device acquires an image frame. That is to say, the frame interval between two adjacent image frames acquired by the image acquisition device and the time interval duration between every two adjacent detection moments may be different.
[0208] Among them, the frame interval between two adjacent image frames acquired by the image acquisition device is determined by the frame rate of the image acquisition device. For example, when the frame rate of the image acquisition device is 50fps (frames per second), the frame interval between two adjacent image frames acquired by the image acquisition device is: 0.02 seconds.
[0209] In one implementation, when reaching a detection moment, if the image acquisition device acquires an image frame at this detection moment, the electronic device can determine that this image frame is the image frame acquired at this detection moment.
[0210] Or, when reaching a detection moment, if the image acquisition device does not acquire an image frame at this detection moment, the electronic device can obtain the most recently acquired image frame by the image acquisition device before this detection moment as the image frame acquired at this detection moment.
[0211] Pixel position representation of the target of interest in the image frame to be processed: The image area occupied by the target of interest in the image frame to be processed. For example, the pixel position of the target of interest in the image frame to be processed can be represented as: . is the pixel coordinate of the upper left corner pixel of the image area occupied by the target of interest in the image frame to be processed, is the width of the image area occupied by the target of interest in the image frame to be processed, is the height of the image area occupied by the target of interest in the image frame to be processed.
[0212] Imaging plane representation: The plane formed by the image coordinate system of the two-dimensional image (i.e., the acquired image frame) formed by the lens in the image acquisition device when acquiring one frame of image. Moreover, the direction represented by the optical axis of the lens is perpendicular to this imaging plane.
[0213] The direction of the imaging plane includes: two mutually perpendicular directions within the imaging plane. For ease of description, the width direction and height direction of the image coordinate system can be respectively determined as the two directions within the imaging plane (i.e., the direction of the imaging plane). Among them, the width direction of the image frame can be denoted as: the direction represented by the x-axis, and the height direction of the image frame can be denoted as: the direction represented by the y-axis.
[0214] In one implementation, the electronic device can obtain the previous detection moment of the first detection moment, and the pixel position of the target of interest in the image frame acquired by the image acquisition device (which can be called the first historical pixel position). Furthermore, calculate the pixel distance between the first historical pixel position and the pixel position of the target of interest in the image frame to be processed.
[0215] It can be understood that the rotation of the lens may cause a fixed position in the real world to have different pixel positions in the image frames captured at different angles. Therefore, in order to determine the influence of the lens rotation on the change of the pixel position of the target of interest, the electronic device can obtain the rotation angle of the lens from the previous detection moment of the first detection moment to the first detection moment. Furthermore, the change amount of the pixel position caused by the rotation angle of the lens can be determined according to the field of view (FOV) of the lens.
[0216] Furthermore, the sum value of the pixel distance between the first historical pixel position and the pixel position of the target of interest in the image frame to be processed, and the change amount of the pixel position caused by the rotation angle of the lens can be obtained, and the ratio of this sum value to the detection time interval is calculated as the moving speed of the target of interest in the direction of the imaging plane at the first detection moment (i.e., the first moving speed).
[0217] Among them, the pixel distance between the first historical pixel position and the pixel position of the target of interest in the image frame to be processed can be: the distance between the pixel coordinates of the center point of the image area represented by the first historical pixel position and the pixel coordinates of the center point of the image area occupied by the target of interest in the image frame to be processed.
[0218] It can be understood that the obtained first moving speed includes: the speed components of the target of interest in two directions in the imaging plane at the third detection moment.
[0219] The moving speed of the target of interest at a detection moment can be denoted as: . represents the speed component of the target of interest in the direction represented by the x-axis at this detection moment; represents the speed component of the target of interest in the direction represented by the y-axis at this detection moment. Among them, The value of can be positive and negative, When the value of is positive, it means that the target of interest moves in the positive direction of the x-axis (the right side in the imaging plane) at this detection moment, When the value of is negative, it means that the target of interest moves in the opposite direction of the x-axis (the left side in the imaging plane) at this detection moment. Similarly, The value of can be positive and negative, When the value of is positive, it means that the target of interest moves in the positive direction of the y-axis (the upper side in the imaging plane) at this detection moment, When the value of is negative, it means that the target of interest moves in the opposite direction of the y-axis (the lower side in the imaging plane) at this detection moment.
[0220] In some embodiments, step S201 includes:
[0221] When the image quality of the image area occupied by the target of interest in this image frame meets the preset conditions, based on the pixel position of the target of interest in this image frame, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed.
[0222] Among them, the preset conditions include at least one of the following: the image area is not overexposed, the image area is not underexposed, and the image area is clear.
[0223] In the embodiments of the present application, after determining the pixel position of the target of interest in the image frame to be processed, the electronic device can perform image quality detection on the image area occupied by the target of interest in the image frame to be processed. Among them, the image area occupied by the target of interest in the image frame to be processed can also be called ROI (Region of Interest).
[0224] In one implementation, the image region occupied by the target of interest in the image frame to be processed is input into a pre-trained image quality detection model to obtain the image quality detection result of the image region occupied by the target of interest in the image frame to be processed.
[0225] Among them, the pre-trained image quality detection model can be: a classification image model trained based on a deep learning network architecture. For example, the deep learning network architecture can be AlexNet (Alex Network), ResNet (Residual Neural Network), etc.
[0226] The image quality detection model is trained based on sample images and corresponding sample labels. The label of a sample image is used to indicate whether the sample image is overexposed, and / or whether the sample image is underexposed, and / or whether the sample image is clear.
[0227] In another implementation, the brightness of each pixel position in the image region occupied by the target of interest in the image frame to be processed can be calculated according to a preset pixel brightness calculation formula. Furthermore, the average value of the brightness of each pixel position in the image region occupied by the target of interest in each image frame to be processed is calculated, and it is determined whether the average value is within a specified brightness interval. If the average value is within the specified brightness interval, it indicates that the image region is not overexposed and the image region is not underexposed. If the average value is greater than the specified brightness interval, it indicates that the image region is overexposed. If the average value is less than the specified brightness interval, it indicates that the image region is underexposed.
[0228] Among them, the brightness of a pixel position can be expressed as: . represents the value of the red channel in the pixel value of this pixel position, represents the value of the green channel in the pixel value of this pixel position, represents the value of the blue channel in the pixel value of this pixel position.
[0229] In yet another implementation, the clarity of the image region occupied by the target of interest in the image frame to be processed can be determined according to a preset clarity evaluation function. For example, the clarity evaluation function can be: the Energy of Gradient (EOG) function.
[0230] Correspondingly, the electronic device can calculate the gray values of the pixel points in the image area occupied by the target of interest in the image frame to be processed. For each pixel point in the image area, calculate the difference in gray values between this pixel point and its adjacent pixel points in the height direction and the width direction respectively. Further, for the height direction, calculate the sum of the squares of the differences in the gray values of the pixel points in the height direction; for the width direction, calculate the sum of the squares of the differences in the gray values of the pixel points in the width direction. And calculate the sum of the sum of squares in the height direction and the sum of squares in the width direction as the energy gradient (EOG) value. If the energy gradient (EOG) value is greater than the preset energy gradient threshold, it indicates that the image area is clear; if the energy gradient (EOG) value is not greater than the preset energy gradient threshold, it indicates that the image area is not clear.
[0231] The preset conditions can be: those that can be preset by technicians according to actual needs. For example, in scenarios with high requirements for image quality, the preset conditions include: the image area is not overexposed, the image area is not underexposed, and the image area is clear.
[0232] Correspondingly, when the image quality of the image area occupied by the target of interest in the image frame to be processed meets the preset conditions, the moving speed (the first moving speed) of the target of interest in the direction of the imaging plane at the first detection moment can be calculated.
[0233] In one implementation, when the image quality of the image area occupied by the target of interest in the image frame to be processed does not meet the preset conditions, the electronic device may not calculate the first moving speed and the moving speed (i.e., the third moving speed) of the target of interest in the direction of the imaging plane at the third detection moment. Correspondingly, when reaching the second detection moment, the electronic device can maintain the rotation speed of the motor in the image acquisition device at the current moment until reaching the third detection moment.
[0234] Based on the above processing, when the image quality of the image area occupied by the target of interest in the image frame to be processed meets the preset conditions, the first moving speed can be calculated. It can be understood that if the image quality of the image area occupied by the target of interest in the image frame to be processed does not meet the preset conditions, it indicates that the quality of the currently acquired image frame to be processed may not be high. Correspondingly, if speed prediction is performed based on this image frame to be processed, it may cause the third moving speed of the target of interest at the third detection moment to be inaccurate, interfering with the subsequent process of adjusting the motor rotation speed. Therefore, only when the image quality of the image area occupied by the target of interest in the image frame to be processed meets the preset conditions, performing the subsequent steps can eliminate the influence of image frames with low image quality on the subsequent control of the image acquisition device and improve the accuracy of the control of the image acquisition device.
[0235] Regarding step S202, the first detection moment, the second detection moment, and the third detection moment are three consecutive adjacent detection moments. Among them, the second detection moment is the next detection moment after the first detection moment, and the third detection moment is the next detection moment after the second detection moment.
[0236] The historical detection moments before the first detection moment can be: the second specified number of consecutive adjacent detection moments before the first detection moment. For example, the second specified number can be 5.
[0237] For each historical detection moment before the first detection moment, the process of obtaining the moving speed (i.e., the second moving speed) of the target of interest in the direction of the imaging plane at this historical detection moment can refer to the above process of calculating the first moving speed, which will not be elaborated here.
[0238] Correspondingly, based on the second moving speed and the first moving speed, prediction can be performed to calculate the moving speed of the target of interest in the direction of the imaging plane at the third detection moment.
[0239] In some embodiments, referring to Figure 3 , Figure 3 is the flowchart of the second type of control after detecting the target of interest provided by the embodiments of the present application. On the basis of Figure 2 , step S202 includes:
[0240] S2021: Determine the pixel distance between the target of interest and the reference stationary object in this image frame.
[0241] S2022: Input the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment, and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed.
[0242] Among them, the speed prediction model is: trained based on sample data and corresponding sample labels.
[0243] The sample data includes: the pixel distance between the sample object and the reference stationary object in the sample image frame collected at a detection moment, the moving speed of the sample object in the direction of the imaging plane at the historical detection moment before this detection moment, and the moving speed of the sample object in the direction of the imaging plane at this detection moment; the sample label corresponding to a sample data includes: the speed of the sample object at the second detection moment after this detection moment.
[0244] In an embodiment of the present application, an electronic device may pre-obtain the pixel positions of a reference stationary object in a to-be-processed image frame. The reference stationary object refers to an object that is stationary in a real-world position. For example, the reference stationary object may include: trees, walls, roadblocks, etc.
[0245] Among them, the process by which the electronic device obtains the pixel positions of the reference stationary object in the to-be-processed image frame may refer to the process of obtaining the pixel positions of the target of interest in the to-be-processed image frame in the above embodiment, which will not be elaborated here.
[0246] Correspondingly, the electronic device may determine the pixel distance between the target of interest and the reference stationary object in the to-be-processed image frame according to the pixel positions of the target of interest in the to-be-processed image frame and the pixel positions of the reference stationary object in the to-be-processed image frame. For example, the distance (such as the Euclidean distance) between the central pixel coordinates included in the pixel position of the target of interest in the to-be-processed image frame and the central pixel coordinates included in the pixel position of the reference stationary object in the to-be-processed image frame may be calculated as the pixel distance between the target of interest and the reference stationary object in the to-be-processed image frame.
[0247] Furthermore, the electronic device may use a pre-trained speed prediction model to predict the moving speed (i.e., the third moving speed) of the target of interest in the direction of the imaging plane at a third detection moment.
[0248] The pre-trained speed prediction model is trained based on a time series network architecture. Among them, the time series network architecture may be: a long short-term memory network (LSTM, Long Short Term Memory Network) structure, an auto-regressive moving average model (ARMA, Auto-Regression and Moving Average Model), etc.
[0249] The speed prediction model is trained based on sample data and corresponding sample labels.
[0250] The sample data includes: the pixel distance between the sample object and the reference stationary object in a sample image frame collected at a detection moment, the moving speed of the sample object in the direction of the imaging plane at a historical detection moment before this detection moment, and the moving speed of the sample object in the direction of the imaging plane at this detection moment; a sample label corresponding to a sample data includes: the speed of the sample object at a second detection moment after this detection moment.
[0251] Furthermore, the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment, and the first moving speed are input into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment, which is used as the third moving speed.
[0252] The second moving speed and the first moving speed can reflect the change amount of the magnitude and direction of the moving speed of the target of interest from the historical detection moment to the first detection moment. Correspondingly, the speed prediction model can combine the change amount of the historical moving speed of the target of interest and the pixel distance between the target of interest and the reference stationary object in the image frame to be processed, and predict the moving speed of the target of interest in the direction of the imaging plane at the third detection moment.
[0253] Based on the above processing, in the process of predicting the moving speed of the target of interest in the direction of the imaging plane at the third detection moment, the electronic device can combine the historical and current moving speeds of the target of interest, and make a prediction based on the pre-trained speed prediction model. During the prediction process, the distance between the target of interest and the reference stationary object in the image frame to be processed is introduced. It can be understood that since the rotation of the lens will cause a fixed position in the real world to present different pixel positions in the image frames at different shooting angles, introducing a reference stationary object can establish a position correlation relationship between the stationary object in the real world and the moving target of interest, thereby providing richer and more stable information for predicting the moving speed of the target of interest, effectively reducing the interference caused by the rotation of the lens, and enhancing the accuracy of speed prediction.
[0254] In another implementation manner, the electronic device can calculate the moving speed of the target of interest in the direction of the imaging plane at the third detection moment, which is used as the third moving speed, based on a preset speed prediction algorithm according to the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment and the first moving speed. The preset speed prediction algorithm can be: the Kalman filter algorithm.
[0255] In yet another implementation manner, the electronic device can input the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment, which is used as the third moving speed.
[0256] Among them, the speed prediction model is trained based on sample data and corresponding sample labels. The sample data includes: the moving speed of the sample object in the direction of the imaging plane at the historical detection moment before a detection moment, and the moving speed of the sample object in the direction of the imaging plane at this detection moment; the sample label corresponding to a sample data includes: the speed of the sample object at the second detection moment after this detection moment.
[0257] The pre-trained speed prediction model is trained based on a time series network architecture. Among them, the time series network architecture can be: a long short-term memory network (LSTM, Long Short Term Memory Network) structure, an autoregressive moving average model (ARMA, Auto-Regression and Moving Average Model), etc.
[0258] Based on the above processing, the electronic device can use the speed prediction model trained by the time series network architecture to predict according to the change trend of the moving speed of the target of interest in the historical time period, and obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment.
[0259] For step S203, when the second detection moment is reached, the electronic device can adjust the rotation speed of the motor according to the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor.
[0260] In some embodiments, see Figure 4 , Figure 4 This is the flowchart of the third type of control after detecting the target of interest provided by the embodiments of the present application. On the basis of Figure 2 , step S203 includes:
[0261] S2031: When the second detection moment is reached, according to the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor, determine the peak rotation speed that the motor needs to reach at the specified moment between the second detection moment and the third detection moment.
[0262] S2032: Use the rotation speed of the motor at the current moment as the initial rotation speed, and adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the peak rotation speed at the specified moment.
[0263] S2033: When the specified moment is reached, if the rotation speed of the motor at the current moment does not reach the maximum rotation speed supported, use the rotation speed of the motor at the current moment as the initial rotation speed, and adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the third moving speed at the third detection moment.
[0264] The method further includes:
[0265] S204: When the specified moment is reached, if the rotational speed of the motor at the current moment is the maximum rotational speed supported, control the motor to maintain the rotational speed at the current moment until the third detection moment.
[0266] In the embodiment of the present application, the specified moment between the second detection moment and the third detection moment is: the moment corresponding to the specified duration after the second detection moment. For example, the specified duration is: a preset multiple of a detection time interval. Among them, the value range of the preset multiple can be [0.5, 0.8]. For example, the preset multiple can be 2 / 3.
[0267] When the second detection moment is reached, the electronic device can determine the peak rotational speed required by the motor at the specified moment between the second detection moment and the third detection moment.
[0268] The electronic device can adjust the rotational speed of the motor in the image acquisition device to adjust the angle (i.e., the direction of the optical axis) of the lens shooting in the image acquisition device, so as to shoot different regions in the real world.
[0269] Among them, the rotational speed of the motor at any moment represents: the magnitude and direction of the angle of rotation of the motor at that moment.
[0270] For example, the rotational speed of the motor at any moment can be recorded as: . represents the rotational speed component of the motor in the direction represented by the x-axis in the imaging plane; represents the rotational speed component of the motor in the direction represented by the y-axis in the imaging plane. Among them, the value of can be positive and negative, the value of being positive means: the motor rotates in the positive direction of the x-axis (the right side in the imaging plane), the value of being negative means: the motor rotates in the opposite direction of the x-axis (the left side in the imaging plane). Similarly, the value of can be positive and negative, the value of being positive means: the motor rotates in the positive direction of the y-axis (the upper side in the imaging plane), the value of being negative means: the motor rotates in the opposite direction of the y-axis (the lower side in the imaging plane).
[0271] It can be understood that as the rotational speed of the motor increases, generally the quality of the image frames acquired by the image acquisition device will decrease. Therefore, technicians can determine the maximum rotational speed supported by the motor according to the relationship between the rotational speed of the motor and the image quality of the acquired image frames, and ensure that when the motor rotates at this maximum rotational speed, the image quality of the acquired image frames meets the preset conditions.
[0272] Regarding step S2032, after determining the peak speed that the motor needs to reach at the specified moment, the electronic device can use the current speed of the motor as the initial speed and adjust the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the peak speed at the specified moment.
[0273] Among them, the electronic device adjusts the speed of the motor by adjusting the pulse frequency. The higher the pulse frequency, the faster the motor rotates; the lower the pulse frequency, the slower the motor rotates.
[0274] In one implementation, the electronic device can calculate the speed difference between the current speed of the motor and the peak speed that the motor needs to reach at the specified moment, and calculate the time difference between the current moment and the specified moment. Furthermore, calculate the ratio of the speed difference to the time difference to obtain the speed acceleration represented by this ratio from the current moment to the specified moment for the motor. Correspondingly, according to the correspondence between the pulse frequency and the motor speed, adjust the pulse frequency in a linear adjustment manner so that the motor maintains the speed acceleration represented by this ratio, which can ensure that the motor reaches the peak speed at the specified moment.
[0275] In another implementation, the electronic device can calculate the speed difference between the current speed of the motor and the peak speed that the motor needs to reach at the specified moment, and calculate the time difference between the current moment and the specified moment. Furthermore, use the sine function to fit the calculated speed difference and time difference so that the electronic device uses a stepped adjustment method to adjust the speed of the motor to make the motor speed approach the sine function. Correspondingly, according to the correspondence between the pulse frequency and the motor speed, adjust the pulse frequency so that the motor accelerates or decelerates according to the fitting result, ensuring that the motor reaches the peak speed at the specified moment. In addition, using the sine function to fit to determine the speed of the motor at each moment from the current moment to the specified moment can gradually adjust the speed of the motor and ensure the smoothness of the motor rotation.
[0276] When the specified moment is reached, if the current speed of the motor does not reach the maximum speed supported, starting from the specified moment, use the current speed of the motor as the initial speed and adjust the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the third moving speed at the third detection moment.
[0277] For example, when the specified moment is reached, if the current speed of the motor is in the same direction as the direction represented by the third moving speed, and the magnitude of the current speed of the motor is greater than the magnitude of the third moving speed, then adjust the speed of the motor in a decreasing manner with the current speed of the motor as the initial speed.
[0278] If the rotation speed of the motor at the current moment is in the same direction as that characterized by the third moving speed, and the magnitude of the rotation speed of the motor at the current moment is less than the magnitude of the third moving speed, then the rotation speed of the motor is adjusted in an increasing manner with the rotation speed of the motor at the current moment as the initial rotation speed.
[0279] If the rotation speed of the motor at the current moment is in the opposite direction to that characterized by the third moving speed, then the magnitude of the rotation speed of the motor is adjusted to 0 in a decreasing manner with the rotation speed of the motor at the current moment as the initial rotation speed; then, with 0 as the initial rotation speed, the rotation speed of the motor is adjusted in an increasing manner according to the magnitude.
[0280] Correspondingly, the rotation speed of the lens in the image acquisition device at the third detection moment is the same as the third moving speed. The rotation speed of the lens is adjusted starting from the second detection moment so that the pixel position change speed of the target of interest in the image frame caused by the rotation of the lens in the image acquisition device at the third detection moment is the same as the pixel position change speed of the target of interest in the image frame caused by the movement of the target of interest itself, which can ensure the synchronization between the rotation speed of the lens in the image acquisition device and the movement speed of the target of interest. That is, the rotation speed of the lens can be adjusted according to the movement speed of the target of interest so that the lens in the image acquisition device can continuously capture the moving target of interest.
[0281] Moreover, from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is the same as the rotation angle of the lens. That is to say, within the time interval from the second detection moment to the third detection moment, the pixel position of the target of interest moving in the image caused by the rotation of the lens can be used to compensate for the pixel position of the target of interest moving in the image caused by the movement of the target of interest itself, ensuring the synchronization between the rotation amount of the lens in the image acquisition device and the movement amount of the target of interest. That is, the movement amount of the target of interest is compensated by the rotation amount of the lens so that the position of the target of interest in the image frame captured within this time interval is stable, and the target of interest can be maintained at the middle position of the image frame captured by the image acquisition device to a certain extent.
[0282] When reaching the specified moment, if the rotation speed of the motor at the current moment is the maximum rotation speed supported, it indicates that when the motor rotates at the maximum rotation speed supported, it may still be unable to enable the image acquisition device to continuously capture the target of interest. At this time, the motor can be controlled to maintain the rotation speed at the current moment until the third detection moment to ensure that the image acquisition device follows the target of interest to the greatest extent.
[0283] It can be understood that if the first detection moment is not the first detection moment during the continuous shooting process of the image acquisition device for the target of interest, then at the previous detection moment of the first detection moment, the electronic device can perform a speed prediction once according to the above steps S201 to S202 to obtain the moving speed of the target of interest in the direction of the imaging plane at the second detection moment. When the first detection moment is reached, the rotation speed of the motor is adjusted. Moreover, the method of adjusting the rotation speed of the motor from the first detection moment to the second detection moment can refer to the method of adjusting the rotation speed of the motor from the second detection moment to the third detection moment as described above, which will not be elaborated here.
[0284] In some embodiments, S2031 includes:
[0285] When the second detection moment is reached, according to a preset formula, based on the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor, determine the peak rotation speed that the motor needs to reach at a specified moment between the second detection moment and the third detection moment.
[0286] Among them, the preset formula is expressed as:
[0287]
[0288] represents the peak rotation speed; is the maximum rotation speed supported by the motor; represents the third moving speed; represents the rotation speed of the motor at the second detection moment.
[0289] In the embodiments of the present application, is preset by those skilled in the art. For example, is 1.5.
[0290] It can be understood that the peak rotation speed that the motor needs to reach at the specified moment includes: the peak rotation speed in the x-axis direction in the imaging plane and the peak rotation speed in the y-axis direction in the imaging plane. [[ID=CO1]] [[ID=CO2]]
[0291] The maximum rotation speed supported by the motor includes: the maximum rotation speed supported by the motor in each direction in the imaging plane. Generally, the rates of the maximum rotation speeds supported by the motor in each direction in the imaging plane are the same.
[0292] Based on the above processing, it is ensured that the peak rotation speed that the motor needs to reach at a specified moment between the second detection moment and the third detection moment is within the maximum rotation speed supported by the motor. In this way, it can ensure the image quality of the image frames collected by the image acquisition device when rotating at the peak rotation speed and ensure the stability of the control of the image acquisition device.
[0293] Taking the determination of the peak rotational speed in the x-axis direction within the imaging plane as an example, and when is 1.5 and the rate of the maximum rotational speed supported by the motor is 7, that is, the rotational speed supported by the motor is [−7, 7]:
[0294] If the rotational speed of the motor at the second detection moment is 3 and the third moving speed is 6, then according to the preset formula, the value of is 7.5, which is greater than the maximum rotational speed 7 supported by the motor. Therefore, the peak rotational speed required by the motor at the specified moment is 7. Correspondingly, when reaching the second detection moment, the electronic device can control the motor to adjust the rotational speed of the motor in an increasing manner from 3 so that the rotational speed of the motor reaches 7 at the specified moment, and control the motor to maintain the rotational speed 7 at the current moment until the third detection moment.
[0295] If the rotational speed of the motor at the second detection moment is 6 and the third moving speed is 3, then according to the preset formula, the value of is 1.5, which is not greater than the maximum rotational speed 7 supported by the motor. Therefore, the peak rotational speed required by the motor at the specified moment is 1.5. Correspondingly, when reaching the second detection moment, the electronic device can control the motor to adjust the rotational speed of the motor in a decreasing manner from 6 so that the rotational speed of the motor reaches 1.5 at the specified moment. When reaching the specified moment, the electronic device can control the motor to adjust the rotational speed of the motor in an increasing manner from 1.5 so that the rotational speed of the motor reaches 3 at the third detection moment.
[0296] If the rotational speed of the motor at the second detection moment is −3 and the third moving speed is 3, then according to the preset formula, the value of is 6, which is not greater than the maximum rotational speed 7 supported by the motor. Therefore, the peak rotational speed required by the motor at the specified moment is 6. Correspondingly, when reaching the second detection moment, the electronic device can control the motor to adjust the rotational speed of the motor in a decreasing manner from −3 to adjust the rotational speed of the motor to 0; then starting from 0 as the initial rotational speed, adjust the rotational speed of the motor in an increasing manner at a rate so that the rotational speed of the motor reaches 6 at the specified moment. When reaching the specified moment, the electronic device can control the motor to adjust the rotational speed of the motor in a decreasing manner from 6 so that the rotational speed of the motor reaches 3 at the third detection moment.
[0297] If the rotational speed of the motor at the second detection moment is 3 and the third moving speed is −3, then according to the preset formula, The value is -6, which is not greater than the maximum rotational speed -7 supported by the motor. Therefore, the peak rotational speed required by the motor at the specified moment is -6. Correspondingly, when reaching the second detection moment, the electronic device can control the motor to adjust the rotational speed of the motor in a decreasing manner from 3 to 0; then, starting from 0 as the initial rotational speed, adjust the rotational speed of the motor in an increasing rate manner so that the rotational speed of the motor reaches -6 at the specified moment. When reaching the specified moment, the electronic device can control the motor to adjust the rotational speed of the motor in a decreasing manner from -6 so that the rotational speed of the motor reaches -3 at the third detection moment.
[0298] The process of determining the peak rotational speed in the y-axis direction within the imaging plane can refer to the process of determining the peak rotational speed in the x-axis direction within the imaging plane, which will not be elaborated here.
[0299] See Figure 5 , Figure 5 which is a schematic flowchart of the control of the image acquisition device provided by the embodiment of the present application. Figure 5 In:
[0300] Target detection means: At the first detection moment, obtain the to-be-processed image frame collected by the image acquisition device, and obtain the pixel positions of each object in the to-be-processed image frame.
[0301] Quality judgment means: The electronic device can detect the image quality of the to-be-processed image frame and / or the image area occupied by each object. When the image quality of the to-be-processed image frame meets the preset conditions, and / or, the image quality of the image area occupied by each object meets the preset conditions, perform subsequent processing.
[0302] Target tracking means: Determine the pixel position of the target of interest in the to-be-processed image frame according to the pixel positions of each object included in the to-be-processed image frame and the pixel positions of the target of interest in the collected historical image frames.
[0303] Velocity prediction means: Predict according to the second moving velocity of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment and the moving velocity of the target of interest in the direction of the imaging plane at the first detection moment (i.e., the first moving velocity), and calculate the moving velocity of the target of interest in the direction of the imaging plane at the third detection moment as the third moving velocity.
[0304] Motor parameter adjustment indication: When the second detection moment is reached, the rotation speed of the motor is adjusted according to the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens.
[0305] It should be noted that the sample images and sample data in this embodiment are all from public data sets.
[0306] Based on the same inventive concept, an embodiment of the present application provides an intelligent image processing device for open-ended custom detection of targets. Refer to Figure 6 , Figure 6 which is a schematic structural diagram of an intelligent image processing device provided by an embodiment of the present application.
[0307] The intelligent image processing device 600 includes:
[0308] A lens 601 for collecting image frames;
[0309] A processor 602 for the above-mentioned intelligent image processing method for open-ended custom detection of targets.
[0310] Based on the intelligent image processing device for open-ended custom detection of targets provided by the embodiment of the present application, it supports the user to input object description information for determining the target of interest in a custom input manner. Correspondingly, based on the object description information, it can be detected whether there is a target of interest in the collected image frame. Furthermore, if a target of interest is detected in the image frame, the image acquisition direction is controlled so that the target of interest is located at a preset position in the image frame. That is to say, the intelligent image processing device for open-ended custom detection of targets can continuously capture the target (i.e., the target of interest) that conforms to the object description information in the image frame, and has high usability.
[0311] In some embodiments, the device further includes: a pan-tilt head;
[0312] The processor is specifically configured to send a control instruction to the pan-tilt head;
[0313] The pan-tilt head is configured to adjust the image acquisition direction of the lens according to the control instruction sent by the processor, so that the target of interest is located at a preset position in the image frame.
[0314] In an embodiment of the present application, the intelligent image processing device includes a pan-tilt head, and a motor for adjusting the image acquisition direction of the lens is disposed within the pan-tilt head. Correspondingly, the pan-tilt head can control the motor to rotate according to a control instruction sent by a processor, so as to adjust the image acquisition direction of the lens, and enable an interested target to be located at a preset position in the image frame.
[0315] Based on the same inventive concept, an embodiment of the present application provides an intelligent image processing system for open-ended custom detection of targets. Refer to Figure 7 , Figure 7 which is a schematic structural diagram of an intelligent image processing system provided by an embodiment of the present application.
[0316] The intelligent image processing system 700 includes:
[0317] A client 701, configured to receive object description information customarily input by a user.
[0318] Among them, the custom input methods include at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in the image frame; the object description information is used to determine the interested target of the user;
[0319] An image processing terminal 702, configured to collect an image frame, and based on the object description information, detect whether there is an interested target in the collected image frame; if an interested target is detected in the image frame, control the image acquisition direction so that the interested target is located at a preset position in the image frame.
[0320] Based on the intelligent image processing system for open-ended custom detection of targets provided by the embodiment of the present application, the client supports the user to input, in a custom input manner, object description information for determining the interested target of the user. Correspondingly, the image processing terminal can detect whether there is an interested target in the collected image frame based on the object description information. Furthermore, if an interested target is detected in the image frame, control the image acquisition direction so that the interested target is located at a preset position in the image frame. That is to say, the electronic device can control the image acquisition device to continuously capture the target (i.e., the interested target) in the image frame that conforms to the object description information according to the actual needs of the user, improving the usability of the image acquisition device.
[0321] In an embodiment of the present application, the client can provide an interaction interface for receiving object description information input by the user, and the client can communicate with the image processing terminal. Correspondingly, the user can input object description information through the interaction interface of the client, and further, the client can send the object description information input by the user to the image processing terminal.
[0322] In one implementation manner, the image processing terminal may be the intelligent image processing device for open and custom detection targets described in the above embodiments.
[0323] In some embodiments, the client is integrated in the image processing terminal.
[0324] In the embodiments of the present application, the client may be integrated inside the image processing terminal. Correspondingly, the image processing terminal may provide an interaction interface for receiving the object description information input by the user to receive the object description information custom-input by the user.
[0325] In some embodiments, the system further includes: a cloud platform;
[0326] The cloud platform is configured to receive the object description information sent by the client and forward it to the image processing terminal.
[0327] In the embodiments of the present application, the client may communicate with the cloud platform. Correspondingly, the client may send the object description information input by the user to the cloud platform. Furthermore, the cloud platform may forward the object description information to the image processing terminal.
[0328] When the performance of the image processing terminal is relatively high, the above steps S102 to S103 may be executed by the image processing terminal. In this way, the computing and communication pressure on the cloud platform can be dispersed, and the efficiency of intelligent image processing can be improved.
[0329] Based on the same inventive concept, the embodiments of the present application provide an intelligent image processing system for open and custom detection targets. Refer to Figure 8 , Figure 8 which is a schematic structural diagram of another intelligent image processing system provided by the embodiments of the present application.
[0330] The intelligent image processing system 800 includes:
[0331] A client 801, configured to receive the object description information custom-input by the user; wherein, the custom-input manner includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in an image frame; the object description information is used to determine the target of interest to the user.
[0332] An image processing terminal 802, configured to collect image frames;
[0333] A cloud platform 803, configured to receive the object description information sent by the client and the image frames sent by the image processing terminal; and based on the object description information, detect whether there is a target of interest in the image frames and send the detection result to the image processing terminal.
[0334] The image processing terminal 802 is further configured to control the direction of image acquisition according to the detection result sent by the cloud platform, so that the target of interest is located at a preset position in the image frame.
[0335] Based on the intelligent image processing system for open custom detection targets provided in the embodiments of the present application, the client supports users to input object description information for determining the target of interest of the user in a custom input manner. Correspondingly, the cloud platform can detect whether there is a target of interest in the acquired image frame based on the object description information. Furthermore, if a target of interest is detected in the image frame, the image processing terminal controls the direction of image acquisition so that the target of interest is located at a preset position in the image frame. That is to say, the electronic device can control the image acquisition device to continuously capture the target (i.e., the target of interest) that conforms to the object description information in the image frame according to the actual needs of the user, improving the usability of the image acquisition device.
[0336] In the embodiments of the present application, the image processing terminal includes a lens and a motor for controlling the direction of image acquisition.
[0337] When the performance of the cloud platform is relatively high, the above steps S102 to S103 can be executed by the cloud platform. In this way, the applicable range of the intelligent image processing method can be ensured, and continuous shooting of the target of interest can be realized when the performance of the image processing terminal is limited.
[0338] Based on the same inventive concept, the embodiments of the present application provide an intelligent image processing device for open custom detection targets. Refer to Figure 9 , Figure 9 which is a structural diagram of an intelligent image processing device for open custom detection targets provided in the embodiments of the present application. The device includes:
[0339] A description information acquisition module 901, configured to acquire object description information input by the user in a custom manner; wherein, the custom input manner includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in the image frame; the object description information is used to determine the target of interest of the user;
[0340] A detection module 902, configured to detect whether there is the target of interest in the acquired image frame based on the object description information;
[0341] A first control module 903, configured to control the direction of image acquisition if the target of interest is detected in the image frame, so that the target of interest is located at a preset position in the image frame.
[0342] In some embodiments, the data format of the object description information is one of text, image, and audio;
[0343] The detection module 902 includes:
[0344] A first detection sub-module, which is used for a captured image frame. If it is detected that the target of interest does not exist in a consecutive specified number of image frames before the current image frame, or the current image frame is the first image frame captured after receiving the object description information, then based on the object description information and the current image frame, use a pre-trained open-vocabulary object detection model to detect whether the target of interest exists in the current image frame;
[0345] A second detection sub-module, which is used for if it is detected that at least one of the consecutive specified number of image frames has the target of interest, then use the open-vocabulary object detection model to detect the current image frame; and determine whether the target of interest exists in the current image frame according to the pixel positions of the detected targets and the pixel positions of the targets of interest detected in the consecutive specified number of image frames.
[0346] In some embodiments, the second detection sub-module is specifically used for:
[0347] Input the object description information and the current image frame into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets in the current image frame that conform to the object description information;
[0348] Or,
[0349] Input the current image frame and a preset prompt word into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets in the current image frame that conform to the preset prompt word.
[0350] In some embodiments, the object description information is triggered by the user in a specified image frame;
[0351] The detection module 902 includes:
[0352] A third detection sub-module, which is used to obtain the pixel positions of each target in the specified image frame by using a pre-trained open-vocabulary object detection model, and determine whether the target of interest exists in the current image frame according to the determined pixel positions of each target and the position information included in the object description information;
[0353] The fourth detection sub-module is used for each image frame after the specified image frame. If at least one image frame among the consecutive specified number of image frames before this image frame is detected to have the target of interest, the pixel positions of each target in this image frame are obtained by using the open-vocabulary object detection model, and based on the determined pixel positions of each target, and the pixel positions of the targets of interest detected in the consecutive specified number of image frames, it is determined whether the target of interest exists in this image frame.
[0354] In some embodiments, the first control module 903 includes:
[0355] The first speed calculation sub-module is used for if the target of interest exists in the image frame obtained at the first detection moment, based on the pixel position of the target of interest in this image frame, calculating the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed;
[0356] The second speed calculation sub-module is used for calculating the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed according to the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment and the first moving speed; wherein, the third detection moment is the next detection moment of the second detection moment; the second detection moment is the next detection moment of the first detection moment;
[0357] The control sub-module is used for when reaching the second detection moment, adjusting the rotation speed of the motor according to the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens.
[0358] In some embodiments, the second speed calculation sub-module is specifically used for:
[0359] Determining the pixel distance between the target of interest and the reference static object in this image frame;
[0360] Inputting the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment, and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed;
[0361] Among them, the speed prediction model is trained based on sample data and corresponding sample labels;
[0362] The sample data includes: the pixel distance between the sample object and the reference stationary object in the sample image frame collected at a detection moment, the moving speed of the sample object in the direction of the imaging plane at a historical detection moment before the detection moment, and the moving speed of the sample object in the direction of the imaging plane at the detection moment; the sample label corresponding to a sample data includes: the speed of the sample object at the second detection moment after the detection moment.
[0363] In some embodiments, the first speed calculation sub-module is specifically configured to:
[0364] When the image quality of the image area occupied by the target of interest in the image frame meets the preset conditions, based on the pixel position of the target of interest in the image frame, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed;
[0365] Among them, the preset conditions include at least one of the following: the image area is not overexposed, the image area is not underexposed, and the image area is clear.
[0366] In some embodiments, the control sub-module is specifically configured to:
[0367] The peak speed determination unit is used to, when reaching the second detection moment, determine the peak speed that the motor needs to reach at a specified moment between the second detection moment and the third detection moment according to the third moving speed, the current speed of the motor in the image acquisition device, and the maximum speed supported by the motor;
[0368] The first speed adjustment unit is used to use the current speed of the motor as the initial speed and adjust the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the peak speed at the specified moment;
[0369] The second speed adjustment unit is used to, when reaching the specified moment, if the current speed of the motor does not reach the maximum speed supported, use the current speed of the motor as the initial speed and adjust the speed of the motor in an increasing or decreasing manner so that the speed of the motor reaches the third moving speed at the third detection moment;
[0370] The device further includes:
[0371] A second control module, configured to, when the specified moment is reached, if the rotation speed of the motor at the current moment is the maximum rotation speed supported by the motor, control the motor to maintain the rotation speed at the current moment until the third detection moment.
[0372] In some embodiments, the peak rotation speed determining unit is specifically configured to:
[0373] When the second detection moment is reached, according to a preset formula, determine the peak rotation speed required for the motor at a specified moment between the second detection moment and the third detection moment based on the third moving speed, the rotation speed of the motor in the image acquisition device at the current moment, and the maximum rotation speed supported by the motor;
[0374] Wherein, the preset formula is expressed as:
[0375]
[0376] represents the peak rotation speed; is the maximum rotation speed supported by the motor; represents the third moving speed; represents the rotation speed of the motor at the second detection moment.
[0377] An embodiment of the present application further provides an electronic device, as Figure 10 shown, including:
[0378] A memory 1001, configured to store a computer program;
[0379] A processor 1002, configured to implement the steps of the intelligent image processing method for an open custom detection target when executing the program stored on the memory 1001. [[ID=A35]]
[0380] And the above-mentioned electronic device may further include a communication bus and / or a communication interface, and the processor 1002, the communication interface, and the memory 1001 complete communication with each other through the communication bus.
[0381] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0382] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0383] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0384] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0385] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned intelligent image processing methods for open-ended custom detection targets are implemented.
[0386] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the above-mentioned intelligent image processing methods for open-ended custom detection targets in the embodiments.
[0387] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0388] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0389] Each embodiment in this specification is described in a related manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the system, device, electronic device, and computer-readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0390] The foregoing are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included within the protection scope of the present application.
Claims
1. An intelligent image processing method for open custom detection targets, characterized in that, The method includes: Obtaining object description information custom - input by the user; wherein, the custom - input method includes at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in an image frame; the object description information is used to determine the user's target of interest. Based on the object description information, detecting whether the target of interest exists in the acquired image frame. If the target of interest is detected in the image frame, controlling the direction of image acquisition so that the target of interest is located at a preset position in the image frame.
2. The method according to claim 1, wherein The data format of the object description information is one of text, image, and audio. The detecting whether the target of interest exists in the acquired image frame based on the object description information includes: For an acquired image frame, if it is detected that the target of interest does not exist in any of the consecutive specified number of image frames before this image frame, or this image frame is the first image frame acquired after receiving the object description information, then based on the object description information and this image frame, using a pre - trained open - vocabulary object detection model to detect whether the target of interest exists in this image frame. If it is detected that at least one of the consecutive specified number of image frames has the target of interest, using the open - vocabulary object detection model to detect this image frame; and based on the pixel positions of the detected targets, and the pixel positions of the targets of interest detected in the consecutive specified number of image frames, determining whether the target of interest exists in this image frame.
3. The method according to claim 2, characterized in that, The using the open - vocabulary object detection model to detect this image frame includes: Inputting the object description information and this image frame into a pre - trained open - vocabulary object detection model to obtain the pixel positions of the targets in this image frame that match the object description information. Or, Inputting this image frame and a preset prompt word into a pre - trained open - vocabulary object detection model to obtain the pixel positions of the targets in this image frame that match the preset prompt word.
4. The method according to claim 1, characterized in that, The object description information is triggered by the user in a specified image frame. The detecting whether the target of interest exists in the acquired image frame based on the object description information includes: Using a pre - trained open - vocabulary object detection model to obtain the pixel positions of each target in the specified image frame, and based on the determined pixel positions of each target and the position information included in the object description information, determining whether the target of interest exists in this image frame. For each image frame after the specified image frame, if it is detected that at least one of the consecutive specified number of image frames before this image frame has the target of interest, using the open - vocabulary object detection model to obtain the pixel positions of each target in this image frame, and based on the determined pixel positions of each target and the pixel positions of the targets of interest detected in the consecutive specified number of image frames, determining whether the target of interest exists in this image frame.
5. The method according to claim 1, wherein If it is detected that the target of interest exists in the image frame, the direction of image acquisition is controlled so that the target of interest is located at a preset position in the image frame, including: If the target of interest exists in the image frame obtained at the first detection moment, based on the pixel position of the target of interest in this image frame, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed; According to the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment and the first moving speed, calculate the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed; wherein, the third detection moment is the next detection moment of the second detection moment; the second detection moment is the next detection moment of the first detection moment; When the second detection moment is reached, adjust the rotation speed of the motor according to the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens.
6. The method according to claim 5, characterized in that, The step of calculating the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed according to the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment and the first moving speed includes: Determine the pixel distance between the target of interest and the reference stationary object in this image frame; Input the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at a historical detection moment before the first detection moment, and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed; Wherein, the speed prediction model is trained based on sample data and corresponding sample labels; The sample data includes: the pixel distance between the sample object and the reference stationary object in the sample image frame collected at a detection moment, the moving speed of the sample object in the direction of the imaging plane at a historical detection moment before this detection moment, and the moving speed of the sample object in the direction of the imaging plane at this detection moment; the sample label corresponding to a sample data includes: the speed of the sample object at the second detection moment after this detection moment.
7. The method according to claim 5, characterized in that, The step of calculating the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed based on the pixel position of the target of interest in this image frame includes: When the image quality of the image area occupied by the target of interest in the image frame meets the preset conditions, based on the pixel positions of the target of interest in the image frame, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment as the first moving speed; Wherein, the preset conditions include at least one of the following: the image area is not overexposed, the image area is not underexposed, and the image area is clear.
8. The method according to claim 5, characterized in that, When the second detection moment is reached, according to the third moving speed, the current rotation speed of the motor in the image acquisition device, and the maximum rotation speed supported by the motor, adjust the rotation speed of the motor so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens, including: When the second detection moment is reached, according to the third moving speed, the current rotation speed of the motor in the image acquisition device, and the maximum rotation speed supported by the motor, determine the peak rotation speed required for the motor at a specified moment between the second detection moment and the third detection moment; Taking the current rotation speed of the motor as the initial rotation speed, adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the peak rotation speed at the specified moment; When the specified moment is reached, if the current rotation speed of the motor does not reach the maximum rotation speed supported, taking the current rotation speed of the motor as the initial rotation speed, adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the third moving speed at the third detection moment; The method further includes: When the specified moment is reached, if the current rotation speed of the motor is the maximum rotation speed supported, control the motor to maintain the current rotation speed until the third detection moment.
9. The method according to claim 8, wherein When the second detection moment is reached, according to the third moving speed, the current rotation speed of the motor in the image acquisition device, and the maximum rotation speed supported by the motor, determine the peak rotation speed required for the motor at a specified moment between the second detection moment and the third detection moment, including: When the second detection moment is reached, according to a preset formula, based on the third moving speed, the current rotation speed of the motor in the image acquisition device, and the maximum rotation speed supported by the motor, determine the peak rotation speed required for the motor at a specified moment between the second detection moment and the third detection moment; Wherein, the preset formula is expressed as: Indicates the peak rotational speed; Is the maximum rotational speed supported by the motor; Indicates the third moving speed; Indicates the rotational speed of the motor at the second detection moment.
10. An intelligent image processing device for an open-type custom detection target, characterized in that, The device includes: A lens for collecting image frames; A processor for executing the method according to any one of claims 1-9.
11. The device according to claim 10, characterized in that, The device further includes: a pan-tilt head; The processor is specifically configured to send a control instruction to the pan-tilt head; The pan-tilt head is used to adjust the image acquisition direction of the lens according to the control instruction sent by the processor, so that the target of interest is located at a preset position in the image frame.
12. An intelligent image processing system for open-ended custom detection targets, characterized in that, The system includes: A client, which is used to receive the object description information custom-input by the user; wherein, the custom-input methods include at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in the image frame; the object description information is used to determine the target of interest of the user. An image processing terminal, which is used to acquire an image frame and, based on the object description information, detect whether the target of interest exists in the acquired image frame; if the target of interest is detected in the image frame, control the image acquisition direction so that the target of interest is located at a preset position in the image frame.
13. The system according to claim 12, wherein The client is integrated in the image processing terminal.
14. The system according to claim 12, wherein The system further includes: a cloud platform; The cloud platform is used to receive the object description information sent by the client and forward it to the image processing terminal.
15. An intelligent image processing system for open-ended custom detection targets, characterized in that, The system includes: A client, which is used to receive the object description information custom-input by the user; wherein, the custom-input methods include at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in the image frame; the object description information is used to determine the target of interest of the user. An image processing terminal, which is used to acquire an image frame; A cloud platform, which is used to receive the object description information sent by the client and the image frame sent by the image processing terminal; and based on the object description information, detect whether the target of interest exists in the image frame and send the detection result to the image processing terminal. The image processing terminal is further used to control the image acquisition direction according to the detection result sent by the cloud platform, so that the target of interest is located at a preset position in the image frame.
16. An intelligent image processing device with an open-ended custom detection target, characterized in that, The device includes: A description information acquisition module, which is used to acquire the object description information custom-input by the user; wherein, the custom-input methods include at least one of the following: inputting text, inputting an image, inputting audio, and triggering an interaction instruction including position information in the image frame; the object description information is used to determine the target of interest of the user. A detection module, which is used to detect whether the target of interest exists in the acquired image frame based on the object description information. A first control module, which is used to control the image acquisition direction if the target of interest is detected in the image frame, so that the target of interest is located at a preset position in the image frame.
17. The device according to claim 16, wherein The data format of the object description information is one of text, image, and audio. The detection module includes: A first detection sub-module, which is used for a single acquired image frame. If it is detected that the target of interest does not exist in a continuous specified number of image frames before this image frame, or this image frame is the first image frame acquired after receiving the object description information, then based on the object description information and this image frame, use a pre-trained open vocabulary object detection model to detect whether the target of interest exists in this image frame. The second detection sub-module is used to, if it detects that at least one of the continuously specified number of image frames has the target of interest, use the open-vocabulary object detection model to detect this image frame; and determine whether the target of interest exists in this image frame according to the pixel positions of the detected targets and the pixel positions of the targets of interest detected in the continuously specified number of image frames; and / or, Specifically, the second detection sub-module is used for: Inputting the object description information and this image frame into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets in this image frame that conform to the object description information; or inputting this image frame and a preset prompt word into a pre-trained open-vocabulary object detection model to obtain the pixel positions of the targets in this image frame that conform to the preset prompt word; and / or, The object description information is triggered by the user in the specified image frame; The detection module includes: The third detection sub-module is used to use a pre-trained open-vocabulary object detection model to obtain the pixel positions of each target in the specified image frame, and determine whether the target of interest exists in this image frame according to the determined pixel positions of each target and the position information included in the object description information; The fourth detection sub-module is used for each image frame after the specified image frame. If it detects that at least one of the continuously specified number of image frames before this image frame has the target of interest, use the open-vocabulary object detection model to obtain the pixel positions of each target in this image frame, and determine whether the target of interest exists in this image frame according to the determined pixel positions of each target and the pixel positions of the targets of interest detected in the continuously specified number of image frames; and / or, The first control module includes: The first speed calculation sub-module is used to, if the target of interest exists in the image frame obtained at the first detection moment, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment based on the pixel position of the target of interest in this image frame, and use it as the first moving speed; The second speed calculation sub-module is used to calculate the moving speed of the target of interest in the direction of the imaging plane at the third detection moment based on the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment and the first moving speed, and use it as the third moving speed; where the third detection moment is the next detection moment after the second detection moment; the second detection moment is the next detection moment after the first detection moment; A control sub-module, configured to, when reaching the second detection moment, adjust the rotation speed of the motor according to the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor, so that the rotation speed of the lens in the image acquisition device at the third detection moment is consistent with the third moving speed, and from the second detection moment to the third detection moment, the moving distance of the target of interest in the direction of the imaging plane is consistent with the rotation angle of the lens; and / or, The second speed calculation sub-module is specifically configured to: Determine the pixel distance between the target of interest and the reference stationary object in the image frame; Input the determined pixel distance, the second moving speed of the target of interest in the direction of the imaging plane at the historical detection moment before the first detection moment, and the first moving speed into a pre-trained speed prediction model to obtain the moving speed of the target of interest in the direction of the imaging plane at the third detection moment as the third moving speed; Wherein, the speed prediction model is trained based on sample data and corresponding sample labels; The sample data includes: the pixel distance between the sample object and the reference stationary object in the sample image frame collected at a detection moment, the moving speed of the sample object in the direction of the imaging plane at the historical detection moment before this detection moment, and the moving speed of the sample object in the direction of the imaging plane at this detection moment; the sample label corresponding to a sample data includes: the speed of the sample object at the second detection moment after this detection moment; and / or, The first speed calculation sub-module is specifically configured to: When the image quality of the image area occupied by the target of interest in the image frame meets a preset condition, calculate the moving speed of the target of interest in the direction of the imaging plane at the first detection moment based on the pixel position of the target of interest in the image frame as the first moving speed; Wherein, the preset condition includes at least one of the following: the image area is not overexposed, the image area is not underexposed, and the image area is clear; and / or, The control sub-module is specifically configured to: A peak rotation speed determination unit, configured to, when reaching the second detection moment, determine the peak rotation speed that the motor needs to reach at a specified moment between the second detection moment and the third detection moment according to the third moving speed, the rotation speed of the motor at the current moment in the image acquisition device, and the maximum rotation speed supported by the motor; A first rotation speed adjustment unit, configured to use the rotation speed of the motor at the current moment as the initial rotation speed and adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the peak rotation speed at the specified moment; A second rotation speed adjustment unit, configured to, when reaching the specified moment, if the rotation speed of the motor at the current moment does not reach the maximum rotation speed supported, use the rotation speed of the motor at the current moment as the initial rotation speed and adjust the rotation speed of the motor in an increasing or decreasing manner so that the rotation speed of the motor reaches the third moving speed at the third detection moment; The device further includes: A second control module, configured to, when the specified moment is reached, if the rotational speed of the motor at the current moment is the maximum rotational speed supported by the motor, control the motor to maintain the rotational speed at the current moment until the third detection moment; And / or The peak rotational speed determination unit is specifically configured to: When the second detection moment is reached, according to a preset formula, determine the peak rotational speed that the motor needs to reach at a specified moment between the second detection moment and the third detection moment based on the third moving speed, the rotational speed of the motor in the image acquisition device at the current moment, and the maximum rotational speed supported by the motor; Wherein, the preset formula is expressed as: Indicates the peak rotational speed; Is the maximum rotational speed supported by the motor; Indicates the third moving speed; Indicates the rotational speed of the motor at the second detection moment.
18. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor, configured to implement the method according to any one of claims 1-9 when executing the program stored on the memory.
19. A computer program product, characterized in that, When the computer program product runs on a computer, cause the computer to execute the method according to any one of claims 1-9.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Vehicle control system based on user input and method thereof
CN107433902A
Video monitoring system and method for target detection and tracking
CN108055501A
Following method, movable platform, following device and storage medium
CN112585944A
A target tracking method, system, device, and computer-readable medium
CN114938429A
Regional target tracking method fusing millimeter wave radar and camera
CN115184917A