PTZ preset point configuration method, electronic device and storage medium
By scene recognition and image division of the initial images collected by the gimbal, and automatically set the gimbal preset point information, the problem of time-consuming and inaccurate manually setting the preset point in the prior art is solved, and efficient and accurate preset point configuration is achieved.
Patent Information
- Application Number
- CN202510327456.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In the existing gimbal cruise plan, the preset point information required for manually setting the gimbal is time-consuming and has low accuracy, resulting in low configuration efficiency.
By obtaining the initial image collected by the target object by the gimbal, scene recognition and image division are performed, and preset point information for each area is automatically set according to the scene category and image division results to which the target object belongs.
It realizes automatic configuration of gimbal preset point information, improves configuration efficiency, and improves the accuracy of preset point information.
Smart Images

Figure CN119851107B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image acquisition technology, and in particular to a pan / tilt preset point configuration method, an electronic device, and a storage medium. Background Art
[0002] The pan-tilt head is an image acquisition device with a rotatable pan-tilt structure. The pan-tilt head is an image acquisition device for acquiring images of targets in different directions. Pan-tilt cruising refers to presetting a number of specific points so that the lens of the pan-tilt head rotates to the corresponding position to obtain a local captured image, thereby achieving the purpose of collecting multiple points, and each point can be collected according to the preset point information in the preset scheme. In the existing pan-tilt cruising scheme, the preset point information required for the pan-tilt head is manually set so that the lens of the pan-tilt head is rotated to the corresponding position to obtain a local captured image. This method of manually setting the preset point information required for the pan-tilt head is time-consuming and has a low accuracy rate.
[0003] In view of the existing technical defects, how to provide a solution that can improve the efficiency of configuring preset point information is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention
[0004] The present application at least provides a pan / tilt preset point configuration method, an electronic device and a storage medium.
[0005] The present application provides a method for configuring preset points of a pan-tilt head, comprising: obtaining an initial image captured by the pan-tilt head of a target object; performing scene recognition on the initial image to obtain a scene category to which the target object belongs; based on the scene category to which the target object belongs, performing image division on the initial image to obtain sub-images corresponding to different areas of the target object, wherein different sub-images contain different areas of the target object; and based on each sub-image, setting preset point information corresponding to each area in the target object.
[0006] The present application provides a device for configuring preset points of a pan-tilt head, comprising: an acquisition module, an identification module, a division module and a setting module; the acquisition module is used to acquire an initial image captured by the pan-tilt head of a target object; the identification module is used to perform scene recognition on the initial image to obtain a scene category to which the target object belongs; the division module is used to perform image division on the initial image based on the scene category to which the target object belongs to obtain sub-images corresponding to different areas of the target object, wherein different sub-images contain different areas of the target object; the setting module is used to set preset point information corresponding to each area in the target object based on each sub-image.
[0007] The present application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-mentioned pan-tilt preset point configuration method.
[0008] The present application provides a computer-readable storage medium on which program instructions are stored. When the program instructions are executed by a processor, the above-mentioned pan-tilt preset point configuration method is implemented.
[0009] The above scheme performs scene recognition on the initial image collected by the pan-tilt head of the target object to obtain the scene category to which the target object belongs, and based on the scene category to which the target object belongs, realizes refined image segmentation of the initial image to obtain sub-images corresponding to different areas of the target object, and different sub-images contain different areas of the target object, so that the sub-images corresponding to each area obtained by the segmentation are more accurate images under the scene category to which the target object belongs, and based on each sub-image, the preset point information corresponding to each area in the target object is set, which can realize automatic configuration of the preset point information corresponding to each area and the configured preset point information is relatively accurate, thereby improving the efficiency of configuring the preset points of the pan-tilt head.
[0010] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.
[0012] Figure 1 This is a schematic diagram of the first process of an embodiment of a method for configuring pan / tilt preset points of the present application;
[0013] Figure 2 This is a second flow diagram of an embodiment of the pan / tilt preset point configuration method of the present application;
[0014] Figure 3 This is a third flow chart of an embodiment of the pan / tilt preset point configuration method of the present application;
[0015] Figure 4a It is a schematic diagram of an initial image in an embodiment of a method for configuring preset points of a pan / tilt platform of the present application;
[0016] Figure 4b This is a first schematic diagram of a sub-image in an embodiment of a method for configuring preset points of a pan / tilt platform of the present application;
[0017] Figure 4c This is a second schematic diagram of a sub-image in an embodiment of the pan / tilt preset point configuration method of the present application;
[0018] Figure 4d It is a third schematic diagram of a sub-image in an embodiment of the pan / tilt preset point configuration method of the present application;
[0019] Figure 5 It is a structural schematic diagram of an embodiment of a pan / tilt preset point configuration device of the present application;
[0020] Figure 6 It is a structural schematic diagram of an embodiment of the electronic device of the present application;
[0021] Figure 7 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0022] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.
[0023] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0024] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the objects associated before and after are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of, for example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C.
[0025] The present application provides some pan-tilt preset point configuration methods and pan-tilt preset point configuration devices. The application scenarios of the pan-tilt preset point configuration method include but are not limited to automatic configuration of pan-tilt preset points, verification of whether the pan-tilt preset point configuration is appropriate, and other scenarios. The executor of the pan-tilt preset point configuration method may be a pan-tilt preset point configuration device. For example, the pan-tilt preset point configuration device may be set in a terminal device or a server or other processing device, wherein the terminal device may be a device for pan-tilt preset point configuration, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, and the like. In some possible implementations, the pan-tilt preset point configuration method may be implemented by a processor calling computer-readable instructions stored in a memory.
[0026] See also Figure 1 , Figure 1 1 is a schematic diagram of a first process of an embodiment of a method for configuring a pan-tilt preset point of the present application. Specifically, the method for configuring a pan-tilt preset point may include the following steps:
[0027] Step S11: Acquire an initial image of the target object captured by the PTZ.
[0028] The pan-tilt head is an image acquisition device with a rotatable pan-tilt head structure. When the pan-tilt head is in different image acquisition scenes, the target object is different. Among them, the image acquisition scene can be a traffic scene, a water conservancy scene, etc. When the image acquisition scene is a traffic scene, the target object can be a road. When the image acquisition scene is a water conservancy scene, the target object can be a river or a road. The target object can be a target acquisition element contained in the image acquisition scene of the pan-tilt head. The target acquisition element is one of several acquisition elements contained in the initial image. It can be understood that the present application takes the image acquisition scene as a traffic scene and the target object as a road in a traffic scene as an example, which will not be repeated below.
[0029] The target object may be a target acquisition element in an image captured by the gimbal. The initial image is an image acquired by the gimbal at any acquisition position. Any acquisition position may be any acquisition angle and / or any acquisition focal length. Exemplarily, the initial image is an image acquired by acquiring an image at any acquisition position when the gimbal rotates according to a preset rule. The preset rule may include an initialization rotation frequency and a number of acquisition positions. Among them, the several acquisition positions may be default acquisition positions, or they may be historical acquisition positions corresponding to historical preset points. The historical preset points are preset points obtained by resetting the preset points last time. The historical acquisition positions are the acquisition angles and acquisition focal lengths of the gimbal for each historical area.
[0030] The above S11 may be that the gimbal rotates according to a preset rule, and performs image acquisition at any acquisition position to which the gimbal is rotated to obtain an initial image corresponding to the arbitrary acquisition position. Specifically, the current acquisition position is determined based on the preset rule. The gimbal performs image acquisition on the target object at the current position to obtain an initial image corresponding to the current position.
[0031] Step S12: Perform scene recognition on the initial image to obtain the scene category to which the target object belongs.
[0032] There is a matching relationship between the image acquisition scene and several preset scene categories of the target object. When the image acquisition scene and the target object are different, the above matching relationship is different. The scene category to which the target object belongs is one of several preset scene categories of the target object. For example, several preset scene categories of the target object include roads such as straight roads, ring roads, one-way roads, two-way roads, auxiliary roads, ramps, tunnels, bridges, overpasses, intersections, T-shaped roads, Y-shaped roads, S-curve roads, spiral roads, etc. After the initial image is subjected to scene recognition, the scene category to which the target object belongs can be a crossroads among the above-mentioned several preset scene categories.
[0033] The above step S12 may be inputting the initial image into a preset scene recognition model to obtain the scene category to which the target object in the initial image belongs output by the preset scene recognition model. The preset scene recognition model may be a model provided with a preset scene recognition algorithm. The preset scene recognition algorithm may be a convolutional neural network, an attention mechanism network, and the like.
[0034] Step S13: based on the scene category to which the target object belongs, the initial image is divided to obtain sub-images corresponding to different regions of the target object.
[0035] Different sub-images contain different regions of the target object. Each region of the target object may be at least part of the image region of the initial image where the target object is located. Each sub-image includes at least one region of the image region where the target object is located, and the region corresponds to the sub-image.
[0036] It can be understood that the initial image includes image regions corresponding to several acquisition elements, and the image regions corresponding to the acquisition elements constitute the initial image. Among them, the target acquisition element, that is, the image region where the target acquisition element is located, is taken as the target image region.
[0037] Image division may be to first divide the target image area in the initial image into several areas, and then divide the initial image according to each area to obtain a sub-image corresponding to each area. Wherein, each sub-image includes at least the area corresponding to the sub-image. Exemplarily, the scene category of the initial image is a crossroad, and the target image area in the initial image is a cross shape, and the cross-shaped target image area is divided into at least four areas. Each area corresponds to a sub-image. Each sub-image includes at least the area corresponding to the sub-image. In some application scenarios, each sub-image includes only one area divided in the target image area corresponding to the target acquisition element and other acquisition elements. In other application scenarios, each sub-image includes one area divided in the target image area corresponding to the target acquisition element, other acquisition elements, and at least part of other areas after division. The other areas after division are other areas in the target image area except the area corresponding to the sub-image.
[0038] The above step S13 may be to determine the target image area in the initial image based on the scene category to which the target object belongs. First, the target image area in the initial image is divided into regions to obtain regions corresponding to the target image area. The initial image is divided based on the regions to obtain sub-images corresponding to the regions. Each sub-image includes at least the region corresponding to the sub-image. Specifically, the above step of determining the target image area in the initial image based on the scene category to which the target object belongs may be to perform image segmentation on the initial image using a preset image segmentation algorithm after determining the scene category to which the target object belongs, to obtain the target image area in the initial image.
[0039] Step S14: Based on each sub-image, preset point information corresponding to each area in the target object is set.
[0040] The preset point information can be expressed as preset point PTZ. Each area corresponds to a preset point.
[0041] The preset point information corresponding to each area includes the acquisition angle of the gimbal for each area, the acquisition focal length of the gimbal for each area, and the acquisition time of the gimbal for each area. Among them, the acquisition angle of each area can be expressed as PT. Specifically, the acquisition angle can be the azimuth of the gimbal in the horizontal and vertical directions respectively obtained by rotating the gimbal up, down, left, and right. The acquisition focal length of each area can be expressed as Z. Specifically, the acquisition focal length can be the final focal length obtained by zooming in and out of the picture at a fixed acquisition angle. The acquisition time of each area is the residence time of the gimbal at each preset point during the cruising process. For example, the residence time of the gimbal at each preset point during the cruising process can be a default time that is equal.
[0042] In some application scenarios, the above step S14 may be to use the acquisition angle and acquisition focal length corresponding to each sub-image as the preset point information of the area corresponding to the sub-image. In other application scenarios, the above step S14 may also be to adjust the acquisition angle and / or acquisition focal length corresponding to each sub-image to obtain the adjusted sub-image. The acquisition angle and / or acquisition focal length of the adjusted sub-image is used as the preset point information of the area corresponding to the sub-image.
[0043] The above scheme performs scene recognition on the initial image collected by the pan-tilt head of the target object to obtain the scene category to which the target object belongs, and based on the scene category to which the target object belongs, realizes refined image segmentation of the initial image to obtain sub-images corresponding to different areas of the target object, and different sub-images contain different areas of the target object, so that the sub-images corresponding to each area obtained by the segmentation are more accurate images under the scene category to which the target object belongs, and based on each sub-image, the preset point information corresponding to each area in the target object is set, which can realize automatic configuration of the preset point information corresponding to each area and the configured preset point information is relatively accurate, thereby improving the efficiency of configuring the preset points of the pan-tilt head.
[0044] In some embodiments, the method for configuring the preset points of the pan-tilt head may further include the following steps: obtaining a preset point reset interval. Determining whether the time interval between the current time and the last preset point reset reaches the preset point reset interval. In response to the time interval between the current time and the last preset point reset reaching the preset point reset interval, executing the step of obtaining an initial image captured by the pan-tilt head of the target object.
[0045] The preset point reset interval is used to indicate the reset time of the preset point information of the gimbal. The preset points of the gimbal can be the above-mentioned areas. In response to the reset time corresponding to the preset point reset interval, the preset points of the gimbal need to be reset. Determine whether the time interval between the current time and the last preset point reset reaches the preset point reset interval. In response to the time interval between the current time and the last preset point reset reaching the preset point reset interval, execute the above step S11 to achieve the next preset point reset. In response to the time interval between the current time and the last preset point reset not reaching the preset point reset interval, the gimbal uses the preset point information of each area obtained after the last preset point reset to perform automatic cruising and collect cruise images corresponding to each area.
[0046] It can be considered that regularly resetting the preset points can ensure that the preset point information required for automatic cruising is relatively accurate preset point information, and can timely correct the preset point information to avoid the gimbal always using wrong preset point information for automatic cruising.
[0047] In some embodiments, the above step S12 may include the following steps: dividing the initial image into image blocks to obtain feature representations of several image blocks in the initial image. Inputting the feature representations of several image blocks into a multi-layer decoder for feature extraction to obtain target features. The multi-layer decoder includes several attention modules arranged in cascade. Classifying the target features to obtain the scene category to which the target object belongs.
[0048] The above-mentioned preset scene recognition model may be provided with an image block division module, a feature extraction module, and a classification processing module. The image block division module may execute the above-mentioned steps of dividing the initial image into image blocks to obtain feature representations of several image blocks in the initial image. Among them, the feature extraction module includes a multi-layer decoder. The multi-layer decoder includes several attention modules arranged in cascade. The feature extraction module may execute the above-mentioned steps of inputting the feature representations of several image blocks into the multi-layer decoder for feature extraction to obtain target features. The classification processing module may execute the above-mentioned steps of classifying the target features to obtain the scene category to which the target object belongs. Among them, a preset feature extraction network may be set on the feature extraction module. Among them, the preset feature extraction network may be a MobileNet feature extraction network, an Inception feature extraction network, etc. The setting method of the feature extraction network is not limited here.
[0049] In some application scenarios, before the feature representations of several image blocks are input into a multi-layer decoder for feature extraction to obtain target features, the method further includes: normalizing the feature representations of several image blocks to obtain normalized feature representations, and using the normalized feature representations as initial feature representations. The method further includes: inputting the initial feature representation into a multi-layer decoder for feature extraction to obtain target features.
[0050] In some application scenarios, the structures of the attention modules are the same but the parameters are different. The features output by the last attention module are used as the features output by the multi-layer decoder, that is, the target features. Each attention module includes a starting attention submodule and a perceptron submodule. The above-mentioned inputting the feature representations of several image blocks into the multi-layer decoder for feature extraction, and obtaining the target features also includes: using the first query key-value pair as the input of the starting attention submodule in the current attention module, obtaining the first feature output by the starting attention submodule in the current attention module, and the query data, key data and value data in the first query key-value pair are respectively obtained by the feature representations of several image blocks or by the features output by the previous attention module, and the current attention module is one of the several attention modules. Specifically, the query data, key data and value data in the first query key-value pair are respectively obtained by the initial feature representation obtained by normalizing the feature representations of several image blocks or by the features output by the previous attention module. The first feature output by the starting attention submodule in the current attention module is fused with the input of the current attention module to obtain the first fused feature. The first fused feature is input into the perceptron submodule in the current attention module for advanced feature extraction to obtain the feature output by the perceptron submodule in the current attention module. The first fused feature is fused with the feature output by the perceptron submodule in the current attention module to obtain a second fused feature, and the second fused feature is used as the feature output by the current attention module.
[0051] In other application scenarios, before the above-mentioned inputting the first fused feature into the perceptron submodule in the current attention module for advanced feature extraction to obtain the feature output by the perceptron submodule in the current attention module, it also includes: normalizing the first fused feature to obtain the normalized first fused feature. The above-mentioned inputting the first fused feature into the perceptron submodule in the current attention module for advanced feature extraction to obtain the feature output by the perceptron submodule in the current attention module also includes: inputting the normalized first fused feature into the perceptron submodule in the current attention module for advanced feature extraction to obtain the feature output by the perceptron submodule in the current attention module. The above-mentioned fusing the first fused feature with the feature output by the perceptron submodule in the current attention module to obtain the second fused feature also includes: fusing the normalized first fused feature with the feature output by the perceptron submodule in the current attention module to obtain the second fused feature.
[0052] Exemplarily, the starting attention submodule may be a Multi-Head Attention module in each attention module, which is used to perform multi-head attention processing on the input features of the starting attention submodule. The perceptron submodule may be an MLP module in each attention module, which is used to perform advanced feature processing on the input features of the perceptron submodule. The above-mentioned normalization processing may be performed using the normalization module in each attention module, and the normalization module may be a Norm module. The above-mentioned fusion processing may be a residual connection processing of the features to be fused to obtain fused features.
[0053] The classification processing module may be a module for setting a preset classification algorithm, and the preset classification algorithm may be a module for mapping the target feature to the scene category to which the target object belongs. The preset classification algorithm may be a fully connected network, a convolutional neural network, a support vector machine, etc. Exemplarily, the target feature is processed by an activation function, and the preset scene category with the largest mapping value among the preset scene categories is taken as the scene category to which the target object belongs.
[0054] Exemplarily, the above-mentioned step of dividing the initial image into image blocks and obtaining the feature representation of several image blocks in the initial image may include: after dividing the input image into image blocks (patches) of fixed size, mapping each image block into a vector representation, and finally all image blocks are transformed into a two-dimensional matrix representation that satisfies the multi-layer decoder input, that is, [num_token, token_dim]. The input initial image is divided into image blocks of the same size, and each image block is usually a non-overlapping square area of fixed size. For example, using ViT-B / 16, the size of the input initial image is 224×224×3 (3 is the RGB channel), and the image is divided into fixed-size patches of size 16×16, then each initial image will generate (224 / 16)^2 = 196 patches, that is, the input sequence length is 196, and the dimension of each patch is 16×16×3=768. It can be understood that the feature representation of several image blocks can be represented as an input sequence.
[0055] For each image block, a linear transformation (convolutional layer) is used to map it into a one-dimensional feature vector. This process is also called Patch Embedding, which converts the pixel information in the image block into a vector of fixed dimension to represent the feature representation of the image block. This is achieved directly through a convolutional layer. When using ViT-B / 16, the convolution kernel size is 16×16, the stride is 16, the number of convolution kernels is 768, and the dimension after linear transformation is 196×768. In addition, a special prediction [class]Token needs to be added at the beginning. This Token is a trainable parameter and its length is also 768, so the final dimension is 197×768.
[0056] In order to avoid the Transformer model not having the ability to perceive spatial information and the image losing its position information after segmentation and rearrangement, it is necessary to add positional encoding to help the feature extraction model understand the position of the image block in the initial image. This parameter is also a trainable parameter. Positional encoding can be understood as a table with N rows in total. The size of N is the same as the length of the input sequence. Each row represents a vector, and the dimension of the vector is the same as the dimension of the input sequence Embedding (768). The operation of positional encoding is addition (sum), not concat. After adding positional encoding information, the dimension is still 197×768.
[0057] The input sequence obtained above is sent to the multi-layer decoder Transformer Encoder for processing. The input features are processed by multiple attention modules Transformer Block, and each block produces a series of feature representations with more semantic information. By repeatedly stacking and superimposing multiple attention modules Transformer Block, the multi-layer decoder can effectively learn the complex features and structural information of the input image. Exemplarily, the multi-layer decoder can be a ViT model.
[0058] The perceptron submodule in the multi-layer decoder can integrate and extract image features to obtain target features, and finally obtain image category information based on the target features. The perceptron submodule can be an MLP Head fully connected feedforward neural network.
[0059] The above-mentioned preset scene recognition model can be a pre-trained scene recognition model. When the image acquisition scene is a traffic scene, different images containing road information (including different road surface materials, objects on different road surfaces, etc.) are input into the above-mentioned preset scene recognition model for training, so that the above-mentioned preset scene recognition model has the ability to recognize roads. It can be understood that when the image acquisition scene is other scenes, for example, the image acquisition scene is a water conservancy scene, different images containing river information (including the color of water in different rivers, the width of rivers on different rivers, etc.) are input into the above-mentioned preset scene recognition model for training, so that the above-mentioned preset scene recognition model has the ability to recognize rivers. In addition to training for roads, the target objects of this application can also be targeted for training for target objects and image acquisition scenes such as rivers and lakes. The model after training can also be used for automatic recognition and automatic cruising of different image acquisition scenes of the gimbal.
[0060] It can be considered that the scene category to which the target object belongs, obtained by dividing the initial image into image blocks, extracting features and classifying them, can improve the accuracy of the scene category to which the target object belongs, thereby improving the accuracy and adaptability of the sub-images obtained by subsequent image division based on the scene category to which the target object belongs, thereby improving the accuracy of the preset point information.
[0061] In some embodiments, the above step S13 may include the following steps: based on the target object scene category, determining a number of turning points of each region, the image region surrounded by the turning points of each region is each region. At least part of each of the turning points is determined as the target turning points of each region. The initial image is cropped according to the position information of the target turning points of each region to obtain a sub-image belonging to the region.
[0062] The turning points of each region are used to indicate that the connecting lines based on the turning points can form each region. The image region surrounded by the turning points of each region is the region. That is, the image region surrounded by the turning points of each region is the target image region corresponding to the region.
[0063] The turning points of each region are regarded as a turning point group. For each region, the key turning points in the turning point group corresponding to the region are regarded as the target turning points of the region. The key turning points in the turning point group are turning points used for image segmentation of the initial image among the turning points. Based on the key turning points, the initial image can be divided into sub-images corresponding to the regions. At least part of the turning points of each region are key turning points among the turning points of the region.
[0064] For each region, the clipping line belonging to the region is determined based on the position information of the target turning point of the region. The initial image is clipped according to the clipping line belonging to the region to obtain the sub-image belonging to the region. The position information of the target turning point is the pixel coordinates of the initial image where the target turning point is located. In some application scenarios, the target turning points of the region are all on the clipping line of the region. In other application scenarios, the target turning points of the region are all on the same side of the clipping line of the region.
[0065] Exemplarily, the pan / tilt rotates to any position, and the image of the position (cut into multiple small images according to the resolution) is sent to the above-mentioned preset scene recognition model to obtain the scene category to which the target object belongs, and the identification information of each image block in the scene category to which the target object belongs. The identification information can indicate whether any image block contains the turning point of the target object. Each image block can be an image block area obtained by pre-dividing the initial image. It can be understood that if any image block in the identification information contains the turning point of the target object, the target image area of the target object in the image block contains intersecting line segments or only intersects with the boundary of the image block. Pre-dividing the initial image by pre-setting the image blocks can be pre-dividing the initial image into image blocks according to the preset division logic so as to record the position coordinates of the initial image where each image block is located. According to the coordinate position of each image block on the initial image, the coordinate information corresponding to several turning points on the initial image can be obtained. If the scene category to which the target object belongs is a complex road (i.e., a T-junction, a crossroads, a three-way intersection), the target image area corresponding to the target object is divided into blocks according to the turning points to obtain each area. If there are overlapping areas in the target image area according to the turning point blocks, the overlapping areas are randomly merged into any area. It can be understood that each area is automatically set as a preset point for subsequent processing steps to determine the preset point information.
[0066] It can be considered that segmenting the initial image according to the scene category to which the target object belongs can improve the accuracy of each sub-image obtained by segmentation.
[0067] See also Figure 2 , Figure 2 This is a second flow chart of an embodiment of a method for configuring pan-tilt preset points of the present application.
[0068] In some embodiments, the preset point information corresponding to each area includes the acquisition angle of the pan / tilt to the area. The above step S14 may include the following steps: for each sub-image, perform the following steps: Figure 2The following steps are shown: Step S21: Obtain the target position information of the region in the sub-image. Step S22: Determine the relative position result of the region based on the target position information. The relative position result is used to indicate the relative relationship between the region and the boundary of the sub-image. Step S23: Determine the acquisition angle of the pan / tilt to the region based on the relative position result of the region.
[0069] The target position information is used to indicate the position information of the sub-image corresponding to the region. The above step S21 may be to redetermine the position information of the region in each sub-image using the coordinate system of the sub-image. It is understandable that the coordinate system of the initial image may be different from the coordinate system of each sub-image. The relative position result is used to indicate the relative relationship between the boundary of the region and the sub-image. Specifically, the relative position result is used to indicate whether the region is in the center of the sub-image corresponding to the region.
[0070] In some application scenarios, the above step S22 may be to determine the center point of the area through the target position information. Whether the distance between the center point of the area and the center point of the sub-image corresponding to the area is less than or equal to the first preset distance. In response to the distance between the center point of the area and the center point of the sub-image corresponding to the area being less than or equal to the first preset distance, the relative position result of the area is determined as the relative relationship between the area and the boundary of the sub-image is a first relative relationship. The first relative relationship is used to indicate that the area is in the center of the sub-image corresponding to the area. In response to the distance between the center point of the area and the center point of the sub-image corresponding to the area being greater than the first preset distance, the relative position result of the area is determined as the relative relationship between the area and the boundary of the sub-image is a second relative relationship. The second relative relationship is used to indicate that the area is not in the center of the sub-image corresponding to the area.
[0071] In other application scenarios, the above step S22 may be to use the target position information to determine the distance between the boundary of the region and the sub-image corresponding to the region, and use the distance as the boundary distance. The boundary distance includes the left boundary distance and / or the right boundary distance between the left boundary and / or the right boundary of the region and the left boundary and / or the right boundary of the sub-image corresponding to the region, and the upper boundary distance and / or the lower boundary distance between the upper boundary and / or the lower boundary of the region and the upper boundary and / or the lower boundary of the sub-image corresponding to the region. Based on the boundary distance, the relative position result of the region is determined. In response to the boundary distance being less than or equal to the second preset distance, the relative position result of the region is determined as the relative relationship between the boundary of the region and the sub-image is the first relative relationship. The first relative relationship is used to indicate that the region is located in the center of the sub-image corresponding to the region. In response to the boundary distance being greater than the second preset distance, the relative position result of the region is determined as the relative relationship between the boundary of the region and the sub-image is the second relative relationship. The second relative relationship is used to indicate that the region is not located in the center of the sub-image corresponding to the region.
[0072] In some application scenarios, the above step S23 may be in response to the relative position result that the area is in the center of the sub-image, determining the acquisition angle and / or acquisition focal length that can acquire the sub-image. Directly use the acquisition angle and / or acquisition focal length that can acquire the sub-image as the preset point information corresponding to the area. Specifically, use the acquisition angle that can acquire the sub-image as the acquisition angle of the pan / tilt to the area. In other application scenarios, the above step S23 may be in response to the relative position result that the area is in the center of the sub-image, judging whether the ratio between the area of the area and the area of the sub-image is within a preset range. In response to the ratio between the area of the area and the area of the sub-image being within a preset range, directly determine the acquisition angle and / or acquisition focal length that can acquire the sub-image. Directly use the acquisition angle and / or acquisition focal length that can acquire the sub-image as the preset point information corresponding to the area. Specifically, the above step S23 at least includes using the acquisition angle that can acquire the sub-image as the acquisition angle of the pan / tilt to the area. In response to the ratio between the area of the region and the area of the sub-image not being within the preset range, the acquisition angle capable of acquiring the sub-image is used as the target acquisition angle. Based on the target acquisition angle, the target acquisition focal length is re-determined. Specifically, based on the target acquisition angle, the target acquisition focal length is re-determined, including: using the acquisition angle capable of acquiring the sub-image as the target acquisition angle, only adjusting the acquisition focal length of the pan / tilt until the ratio between the area of the adjusted region obtained by adjusting the acquisition focal length and the area of the adjusted sub-image is within the preset range, and using the adjusted acquisition focal length of the adjusted sub-image as the target acquisition focal length. It can be understood that, according to different acquisition focal lengths, the region in the adjusted sub-image will be enlarged or reduced, and the region obtained by adjusting the acquisition focal length is the adjusted region, and the adjusted region is different from the region corresponding to the acquisition focal length before adjustment. Directly use the target acquisition angle and / or the target acquisition focal length as the preset point information corresponding to the region. Specifically, the above step S23 at least includes using the target acquisition angle as the acquisition angle of the pan / tilt for the region.
[0073] In other application scenarios, the above step S23 may be to determine the target acquisition angle in response to the above relative position result that the area is not in the center of the sub-image. Determining the target acquisition angle includes: only adjusting the acquisition angle of the gimbal until the adjusted area obtained by the adjusted acquisition angle is in the center of the adjusted sub-image, and using the adjusted acquisition angle of the adjusted sub-image as the target acquisition angle. Specifically, the above step S23 at least includes using the target acquisition angle as the acquisition angle of the gimbal for the area.
[0074] It can be considered that, by determining the acquisition angle of the pan / tilt platform for each area according to the relative position results of each area, the accuracy of the preset point information corresponding to each area can be improved.
[0075] In some embodiments, the relative position result includes that the region is in the center of the sub-image, or the region is not in the center of the sub-image. The above step S23 may include the following steps: in response to the relative position result that the region is not in the center of the sub-image, obtain the first image acquired by the gimbal at the target angle to capture the region and use the first image as the first advanced image. The region is in the center of the first image. Or, in response to the relative position result that the region is in the center of the sub-image, use the sub-image as the first advanced image. Then, based on the first advanced image, determine the acquisition angle of the gimbal to the region.
[0076] The target angle is used to indicate that the adjusted area obtained by adjusting the acquisition angle of the PTZ until the adjusted area is in the center of the adjusted sub-image, and the acquisition angle of the adjusted sub-image can be acquired. The adjusted area is the image area related to the above area in the adjusted sub-image obtained by adjusting the acquisition angle of the PTZ until the adjusted area is adjusted.
[0077] Based on the first advanced image, determining the acquisition angle of the gimbal to the area may be to use the acquisition angle and / or acquisition focal length that can acquire the first advanced image as the preset point information of the gimbal to the area. Specifically, the above step of determining the acquisition angle of the gimbal to the area based on the first advanced image at least includes using the acquisition angle that can acquire the first advanced image as the acquisition angle of the gimbal to the area.
[0078] Exemplarily, determine whether the coordinate position of the area in the sub-image is in the center of the coordinate system of the entire screen of the sub-image. Specifically, determine whether the values of the left and right margins and the upper and lower margins of each area are similar. Among them, the left and right margins of each area are the distances between the left and right boundaries of the area and the left and right boundaries of the sub-image. The upper and lower margins of each area are the distances between the upper and lower boundaries of the area and the upper and lower boundaries of the sub-image. If the values of the left and right margins and the upper and lower margins of each area are similar, the acquisition angle of the sub-image can be collected as the acquisition angle of the pan-tilt head to the area. If the values of the left and right margins and the upper and lower margins of each area are not similar, the PT position of the area is automatically adjusted so that the adjusted area corresponding to the area is displayed in the center of the screen of the adjusted sub-image corresponding to the area.
[0079] It can be considered that determining the acquisition angle of the pan / tilt platform for each area according to whether each area is located in the center of the sub-image corresponding to each area can improve the accuracy of the preset point information corresponding to each area.
[0080] See also Figure 3 , Figure 3 This is a third flow chart of an embodiment of a method for configuring pan-tilt preset points of the present application.
[0081] In some embodiments, the preset point information corresponding to each area includes the acquisition focal length of the pan / tilt to the area. The above step S14 may include the following steps: for each sub-image, perform the following Figure 3 The following steps are shown: Step S31: The area occupied by the region in the sub-image is used as the initial area. Step S32: In response to the initial area being within the preset area range, the sub-image is used as the second advanced image. Or, Step S33: In response to the initial area not being within the preset area range, a second image acquired by the gimbal at the target focal length of the region is acquired and the second image is used as the second advanced image. Among them, the area occupied by the region in the second image is within the preset area range. Step S34: Based on the second advanced image, the gimbal is set to acquire the focal length of the region.
[0082] The initial area is the area of any area in the sub-image corresponding to the area. The acquisition angle that can acquire the sub-image is used as the target acquisition angle, and the acquisition focal length of the gimbal is adjusted until the area of the adjusted area obtained by adjusting the acquisition focal length is within the preset area range, and the adjusted acquisition focal length of the adjusted sub-image is used as the target acquisition focal length. The second image is the above-mentioned adjusted sub-image.
[0083] The step of setting the acquisition focal length of the pan / tilt head for the area based on the second advanced image may be to use the acquisition angle and / or acquisition focal length that can acquire the second advanced image as the preset point information of the pan / tilt head for the area. Specifically, the step of setting the acquisition focal length of the pan / tilt head for the area based on the second advanced image at least includes using the acquisition focal length that can acquire the second advanced image as the acquisition focal length of the pan / tilt head for the area.
[0084] Exemplarily, the relationship between the area size M of the sub-image corresponding to the area occupied by the judgment area and the preset area range, the maximum value in the preset area range is the upper limit value of the initial area, the minimum value in the preset area range is the lower limit value of the initial area, and the preset area range can be set to one-third to one-half of the overall area of the sub-image. If the initial area is greater than the upper limit value Mmax of the preset area range, then the lens magnification Z is reduced, that is, the acquisition focal length is reduced, and the calculation is repeated until the adjusted area is at least equal to the upper limit value of the preset area range. The adjusted area is the area of the adjusted sub-image occupied by the adjusted area. If the initial area is less than the lower limit value Mmin of the preset area range, then the lens magnification Z is increased, that is, the acquisition focal length is increased, and the calculation is repeated until the adjusted area is at least equal to the lower limit value of the preset area range. The adjusted area is the area of the adjusted sub-image occupied by the adjusted area.
[0085] In some application scenarios, before executing the above step S31, it is first determined whether the relative position result is that the region is in the center of the sub-image. If the determination result is that the relative position result is that the region is in the center of the sub-image, the above step S31 is executed. In other application scenarios, before executing the above step S31, it is first determined whether the relative position result is that the region is in the center of the sub-image. In response to the determination result that the relative position result is that the region is not in the center of the sub-image, the acquisition angle of the region is first adjusted until the adjusted sub-image acquired by the adjusted acquisition angle makes the adjusted region in the adjusted sub-image in the center of the adjusted sub-image. The adjusted acquisition angle is used as the target acquisition angle, and the above step S31 is executed. It can be considered that after determining the target acquisition angle first, determining the target acquisition focal length according to the target acquisition angle can improve the accuracy of determining the preset point information of the region. Specifically, the PTZ value (that is, the target acquisition angle and the target acquisition focal length) finally obtained for the region is automatically set as the final preset point coordinate, and the preset point information setting of the next region is performed until all regions in the screen of the initial image are set as detection preset points. Then, the above step S11 is executed again to set the preset points of all areas in the next initial image until the pan / tilt rotates according to the preset rule and all the initial images have the preset points set. The fact that all the initial images have the preset points set indicates that the current preset points have been reset, so that the pan / tilt performs automatic cruising according to the preset point information after the current preset points are reset.
[0086] It can be considered that determining the acquisition focal length of the pan / tilt platform for each area according to whether the initial area of each area is within the preset area range can improve the accuracy of the preset point information corresponding to each area.
[0087] In some embodiments, the preset point information corresponding to each area includes the acquisition time of the pan / tilt system for the area, and the above step S14 may include the following steps: using the area of the initial image occupied by each area as the reference area of each area. Using the ratio between the reference areas of each area as the reference ratio. Based on the reference ratio, setting the acquisition time of the pan / tilt system for the area.
[0088] The base area of each region is used to represent the area of the initial image occupied by the region. The base ratio is the ratio between the base areas of each region.
[0089] In some application scenarios, the acquisition time corresponding to the initial image acquired in the above step S11 is a default time. According to the reference ratio of the initial image and the default time, the acquisition time is allocated to each area obtained by dividing the initial image.
[0090] Exemplarily, the dwell time of the preset point corresponding to each area during the cruising process of the gimbal can be set to a fixed dwell time, or the acquisition time of each area can be dynamically allocated according to the reference ratio. For example, the target image area in the initial image is divided into 3 areas. The reference ratio is 1:1:3. The default time is 5 minutes. According to the reference ratio and default time of the initial image, the acquisition times allocated to each area divided by the initial image are 1 minute, 1 minute, and 3 minutes, respectively. The gimbal automatically cruises according to the acquisition time of each area. It can be considered that determining the acquisition time of the gimbal for each area according to the reference ratio between the areas can improve the accuracy of the preset point information corresponding to each area determined.
[0091] See also Figure 4a , Figure 4a It is a schematic diagram of the initial image in an embodiment of the pan-tilt preset point configuration method of the present application.
[0092] like Figure 4a In the initial image shown in FIG. 1 , the target object is a road in a traffic scene. The scene category to which the target object belongs is a crossroad. The target image area is the image area surrounded by turning points P1 to P12. The regions of the initial image determined in the above step S13 are respectively as follows: Figure 4a Region 1, Region 2, Region 3 and Region 4 are shown.
[0093] See also Figure 4b , Figure 4c as well as Figure 4d , Figure 4b This is a first schematic diagram of a sub-image in an embodiment of a method for configuring preset points of a pan / tilt platform of the present application. Figure 4c This is a second schematic diagram of a sub-image in an embodiment of a method for configuring pan / tilt preset points of the present application. Figure 4d This is the third schematic diagram of the sub-image in the first embodiment of the pan / tilt preset point configuration method of the present application.
[0094] like Figure 4b As shown, line segment a is the above-mentioned cropping line. After performing the above-mentioned step of dividing the initial image based on the scene category to which the target object belongs to obtain sub-images corresponding to different regions of the target object, sub-image 1 is the sub-image corresponding to the above-mentioned region 2. It can be understood that different sub-images contain different regions of the target object. For example, the sub-image corresponding to region 2 includes at least Figure 4a The image region where region 2 is located in the initial image shown in FIG. The sub-image corresponding to region 2 includes Figure 4a The image area where area 2 is located in the initial image shown, as well as at least a portion of area 1 and at least a portion of area 3. Figure 4b The sub-image 1 shown is a sub-image corresponding to the area 2 obtained after executing the above step S13.
[0095] The above step S21 can be to obtain the area 2 in such a way that Figure 4b In response to the relative position result that the area is not in the center of the sub-image, a first image acquired by the gimbal at the target angle for the area is obtained and the first image is used as the first advanced image. Figure 4c The sub-image 2 shown in FIG. The acquisition angle that can acquire the first advanced image is used as the target acquisition angle. The target acquisition angle is used as the acquisition angle in the preset point information corresponding to area 2.
[0096] The above step S31 may be as follows Figure 4c The area occupied by the region 2 in the sub-image 2 shown is taken as the initial area. The above step S33 may be that in response to the initial area not being within the preset area range, a second image acquired by the gimbal at the target focal length of the region is acquired and the second image is taken as the second advanced image. The second advanced image is as follows Figure 4d The sub-image 3 shown. The acquisition angle and acquisition focal length that can acquire the second advanced image are respectively used as the target acquisition angle and target acquisition focal length. The target acquisition angle and target acquisition focal length are respectively used as the acquisition angle and acquisition focal length in the preset point information corresponding to area 2.
[0097] It can be considered that the present application can automatically adjust the PTZ value according to the scene category to which the target object belongs, that is, automatically determine the preset point information corresponding to each area, so that the image captured by the PTZ meets the detection requirements, and the target objects in different environmental scenes can be automatically identified for monitoring without manual intervention configuration. At the same time, the first preset distance and the second preset distance and the preset area range in the image boundary data provided by the present application make the cruise more targeted and automatically exclude areas that do not need to be detected.
[0098] The above scheme performs scene recognition on the initial image collected by the pan-tilt head of the target object to obtain the scene category to which the target object belongs, and based on the scene category to which the target object belongs, realizes refined image segmentation of the initial image to obtain sub-images corresponding to different areas of the target object, and different sub-images contain different areas of the target object, so that the sub-images corresponding to each area obtained by the segmentation are more accurate images under the scene category to which the target object belongs, and based on each sub-image, the preset point information corresponding to each area in the target object is set, which can realize automatic configuration of the preset point information corresponding to each area and the configured preset point information is relatively accurate, thereby improving the efficiency of configuring the preset points of the pan-tilt head.
[0099] See also Figure 5 , Figure 5It is a structural schematic diagram of an embodiment of the pan-tilt preset point configuration device of the present application. The pan-tilt preset point configuration device 50 includes an acquisition module 51, an identification module 52, a division module 53 and a setting module 54; the acquisition module 51 is used to acquire the initial image captured by the pan-tilt of the target object; the identification module 52 is used to perform scene recognition on the initial image to obtain the scene category to which the target object belongs; the division module 53 is used to perform image division on the initial image based on the scene category to which the target object belongs, and obtain sub-images corresponding to different areas of the target object, and different sub-images contain different areas of the target object; the setting module 54 is used to set the preset point information corresponding to each area in the target object based on each sub-image.
[0100] The above scheme performs scene recognition on the initial image collected by the pan-tilt head of the target object to obtain the scene category to which the target object belongs, and based on the scene category to which the target object belongs, realizes refined image segmentation of the initial image to obtain sub-images corresponding to different areas of the target object, and different sub-images contain different areas of the target object, so that the sub-images corresponding to each area obtained by the segmentation are more accurate images under the scene category to which the target object belongs, and based on each sub-image, the preset point information corresponding to each area in the target object is set, which can realize automatic configuration of the preset point information corresponding to each area and the configured preset point information is relatively accurate, thereby improving the efficiency of configuring the preset points of the pan-tilt head.
[0101] For the functions performed by each module, please refer to the PTZ preset point configuration method, which will not be repeated here.
[0102] See also Figure 6 , Figure 6 6 is a schematic diagram of the structure of an embodiment of an electronic device of the present application. The electronic device 60 includes a memory 61 and a processor 62, and the processor 62 is used to execute the program instructions stored in the memory 61 to implement the steps in the above-mentioned gimbal preset point configuration method embodiment. In a specific implementation scenario, the electronic device 60 may include but is not limited to: a multi-camera device, a microcomputer, and a server. In addition, the electronic device 60 may also include a laptop computer, a tablet computer and other mobile devices, which are not limited here.
[0103] Specifically, the processor 62 is used to control itself and the memory 61 to implement the steps in the above-mentioned PTZ preset point configuration method embodiment. The processor 62 can also be called a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 62 can be implemented by an integrated circuit chip.
[0104] The above scheme performs scene recognition on the initial image collected by the pan-tilt head of the target object to obtain the scene category to which the target object belongs, and based on the scene category to which the target object belongs, realizes refined image segmentation of the initial image to obtain sub-images corresponding to different areas of the target object, and different sub-images contain different areas of the target object, so that the sub-images corresponding to each area obtained by the segmentation are more accurate images under the scene category to which the target object belongs, and based on each sub-image, the preset point information corresponding to each area in the target object is set, which can realize automatic configuration of the preset point information corresponding to each area and the configured preset point information is relatively accurate, thereby improving the efficiency of configuring the preset points of the pan-tilt head.
[0105] See also Figure 7 , Figure 7 The computer-readable storage medium 70 stores program instructions 701, which, when executed by a processor, implement the steps in any of the above-mentioned PTZ preset point configuration method embodiments.
[0106] The above scheme performs scene recognition on the initial image collected by the pan-tilt head of the target object to obtain the scene category to which the target object belongs, and based on the scene category to which the target object belongs, realizes refined image segmentation of the initial image to obtain sub-images corresponding to different areas of the target object, and different sub-images contain different areas of the target object, so that the sub-images corresponding to each area obtained by the segmentation are more accurate images under the scene category to which the target object belongs, and based on each sub-image, the preset point information corresponding to each area in the target object is set, which can realize automatic configuration of the preset point information corresponding to each area and the configured preset point information is relatively accurate, thereby improving the efficiency of configuring the preset points of the pan-tilt head.
[0107] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0108] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0109] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0110] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0111] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
Claims
1. A method for configuring pan / tilt preset points, characterized in that: The method comprises: Obtaining the initial image of the target object captured by the PTZ; Performing scene recognition on the initial image to obtain the scene category to which the target object belongs; Based on the scene category to which the target object belongs, the initial image is segmented to obtain sub-images corresponding to different regions of the target object, wherein different sub-images contain different regions of the target object; Based on each of the sub-images, preset point information corresponding to each of the regions in the target object is set, wherein the preset point information corresponding to each of the regions includes a focal length of the pan / tilt head for capturing the region. The setting of preset point information corresponding to each area in the target object based on each sub-image includes: performing the following steps for each sub-image: taking the area occupied by the area in the sub-image as the initial area; in response to the initial area being within a preset area range, taking the sub-image as a second advanced image; or, in response to the initial area not being within the preset area range, acquiring a second image of the area captured by the gimbal at a target focal length and taking the second image as a second advanced image, wherein the area occupied by the area in the second image is within the preset area range; and setting the capture focal length of the gimbal for the area based on the second advanced image.
2. The method according to claim 1, characterized in that The preset point information corresponding to each of the areas includes the acquisition angle of the pan / tilt platform to the area, and the setting of the preset point information corresponding to each of the areas in the target object based on each of the sub-images includes: For each of the sub-images, perform the following steps: Acquire target position information of the area in the sub-image; Based on the target position information, determining a relative position result of the region, wherein the relative position result is used to represent a relative relationship between the region and a boundary of the sub-image; According to the relative position result of the area, the acquisition angle of the pan / tilt head for the area is determined.
3. The method according to claim 2, characterized in that The relative position result includes that the area is located at the center of the sub-image, or the area is not located at the center of the sub-image, and determining the acquisition angle of the pan / tilt head to the area according to the relative position result of the area includes: In response to the relative position result indicating that the region is not located at the center of the sub-image, acquiring a first image acquired by the gimbal at a target angle for the region and using the first image as a first advanced image, wherein the region is located at the center of the first image; or In response to the relative position result that the region is located at the center of the sub-image, using the sub-image as a first advanced image; Based on the first advanced image, a capture angle of the pan / tilt head for the area is determined.
4. The method according to any one of claims 1 to 3, characterized in that The performing scene recognition on the initial image to obtain the scene category to which the target object belongs includes: Dividing the initial image into image blocks to obtain feature representations of a plurality of image blocks in the initial image; Inputting the feature representations of the plurality of image blocks into a multi-layer decoder for feature extraction to obtain target features, wherein the multi-layer decoder includes a plurality of attention modules arranged in cascade; The target features are classified to obtain the scene category to which the target object belongs.
5. The method according to any one of claims 1 to 3, characterized in that: The preset point information corresponding to each of the areas includes the acquisition time of the area by the pan / tilt platform, and the setting of the preset point information corresponding to each of the areas in the target object based on each of the sub-images includes: Taking the area of the initial image occupied by each of the regions as the reference area of each of the regions; The ratio between the base areas of the said regions is taken as the base ratio; Based on the reference ratio, the acquisition time of the pan / tilt platform for the area is set.
6. The method according to any one of claims 1 to 3, characterized in that: The step of dividing the initial image based on the scene category to which the target object belongs to obtain sub-images corresponding to different regions of the target object includes: Based on the scene category of the target object, determining a number of turning points of each of the regions, where the image region surrounded by the number of turning points of each of the regions is each of the regions; determining at least part of each of the plurality of turning points as target turning points of each of the regions; The initial image is cropped according to the position information of the target turning point of each of the regions to obtain a sub-image to which the region belongs.
7. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Get the preset point reset interval; Determine whether the time interval between the current time and the last preset point reset reaches the preset point reset interval; In response to the time interval between the current time and the last preset point reset reaching the preset point reset interval, the step of acquiring the initial image captured by the pan / tilt head of the target object is performed.
8. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Preset position adjusting method and device based on pan-tilt camera
CN114040094A
Method, device and equipment for analysis area determination and intelligent analysis
CN118397255A