Video monitoring holder control method and system and electronic equipment
Through depth estimation and gesture recognition models, pan-tilt control based on gesture images is realized, which solves the problems of complexity and inconvenience of traditional control methods, improves the convenience and flexibility of pan-tilt control, and enhances the accuracy of gesture recognition.
Patent Information
- Application Number
- CN202510530726.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-09-19
AI Technical Summary
Existing video surveillance PTZ control methods rely on physical buttons, remote controls, or software interface operations, which are complex, inconvenient, and lack flexibility. This reduces the convenience and flexibility of PTZ control, especially in scenarios that require flexible mobile operations.
By acquiring the user's gesture image, gesture recognition is performed using the depth estimation model and gesture recognition model to identify the start control gesture, end control gesture and gimbal positioning gesture, and generate corresponding control instructions to achieve the rotation and positioning of the gimbal.
The convenience and flexibility of PTZ control are improved. Operators can accurately control the PTZ remotely with simple gestures, which enhances the accuracy of gesture recognition and the ability to adapt to complex scenarios.
Smart Images

Figure CN120669847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video surveillance technology, and in particular to a video surveillance pan / tilt control method, system and electronic equipment. Background Art
[0002] In the fields of security monitoring, photography and videography, pan / tilt heads are widely used to adjust the shooting angles of cameras and other equipment.
[0003] Traditional PTZ control methods rely primarily on physical buttons, remote controls, or software interfaces. Physical button operation is not convenient enough, requiring operators to be in close proximity to the device, making it difficult to quickly adjust the PTZ angle in complex scenarios. While remote controls can achieve control from a distance, they are inconvenient to carry and easy to lose. Furthermore, the numerous buttons on the remote control make operation complex, making it difficult for unfamiliar users to quickly and accurately perform the required operations. Control via a software interface requires operators to use devices such as computers, especially in scenarios requiring flexible mobile operations. This approach has significant limitations, reducing the convenience and flexibility of video surveillance PTZ control. Summary of the Invention
[0004] The present invention provides a video surveillance pan-tilt control method, system and electronic equipment to solve the problem in the prior art that control through a software interface requires an operator to use a computer or other equipment, especially in scenarios requiring flexible mobile operation. This method has great limitations and reduces the convenience and flexibility of video surveillance pan-tilt control.
[0005] The present invention provides a video surveillance PTZ control method, comprising the following steps: Get the user's gesture image; Inputting the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; Determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result includes a start control gesture, an end control gesture, and a gimbal positioning gesture; the start control gesture is used to activate a gimbal control subsystem, the gimbal positioning gesture is used to control the gimbal rotation of the gimbal control subsystem, and the end control gesture is used to stop the gimbal rotation; Generate a control instruction corresponding to the gesture recognition result.
[0006] According to a video surveillance pan-tilt control method provided by the present invention, determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image includes: Inputting the gesture depth information and the gesture image into a gesture recognition model to obtain a gesture recognition result corresponding to the gesture image output by the gesture recognition model; The gesture recognition model includes a three-dimensional key point sequence extraction module, a gesture feature extraction module, a fusion module and a classifier; The three-dimensional key point sequence extraction module is used to extract the time sequence features corresponding to the three-dimensional key point sequence in the gesture depth information; The gesture feature extraction module is used to extract spatial features in the gesture image; The fusion module is used to determine a fusion feature based on the temporal feature and the spatial feature; The classifier is used to determine a gesture recognition result corresponding to the gesture image based on the fusion feature.
[0007] According to a video surveillance pan-tilt control method provided by the present invention, the gesture recognition result includes the coordinates of the finger key points and the gesture classification result; Inputting the gesture depth information and the gesture image into a gesture recognition model to obtain a gesture recognition result corresponding to the gesture image output by the gesture recognition model includes: The gesture depth information and the gesture image are input into a gesture recognition model to obtain the gesture key point coordinates and the gesture classification result output by the gesture recognition model.
[0008] According to a video surveillance pan-tilt control method provided by the present invention, the gesture depth information and the gesture image are input into a gesture recognition model to obtain the gesture key point coordinates and the gesture classification result output by the gesture recognition model, and then the method further includes: Based on the coordinates of the finger key points, the direction and angle of each finger are determined.
[0009] According to a video surveillance pan-tilt control method provided by the present invention, determining the direction and angle of each finger based on the coordinates of the finger key points includes: Determining the direction pointed by each finger based on the fingertip coordinates and the finger base coordinates in the finger key point coordinates; Determine a first vector based on the middle finger joint coordinates and the fingertip coordinates in the finger key point coordinates, and determine a second vector based on the middle finger joint coordinates and the finger base coordinates; The angle pointed by each finger is determined based on the angle between the first vector and the second vector.
[0010] According to a video surveillance PTZ control method provided by the present invention, the training step of the gesture recognition model includes: Acquire an initial gesture recognition model; the initial gesture recognition model includes an initial three-dimensional key point sequence extraction module, an initial gesture feature extraction module, an initial fusion module, a regression branch, and a classification branch; Acquire a sample gesture image and a labeled gesture classification result of the sample gesture image; the labeled gesture classification result includes the labeled finger key point coordinates and the labeled gesture classification result; determining sample gesture depth information of the sample gesture image; Extracting sample time sequence features corresponding to the three-dimensional key point sequence in the sample gesture depth information based on the initial three-dimensional key point sequence extraction module; Extracting sample spatial features from the sample gesture image based on the gesture feature extraction module; Performing feature fusion on the temporal features and the spatial features based on the fusion module to obtain a sample fusion feature; Predicting the coordinates of the finger key points on the sample fusion features based on the regression branch to obtain the predicted coordinates of the finger key points; Performing gesture classification prediction on the sample fusion features based on the classification branch to obtain a predicted gesture classification result; Based on the predicted gesture classification result and the labeled gesture classification result, as well as the predicted finger key point coordinates and the labeled finger key point coordinates, a target loss is determined, and parameters of the initial gesture recognition model are iterated based on the target loss to obtain the gesture recognition model.
[0011] According to a video surveillance PTZ control method provided by the present invention, the step of obtaining a user's gesture image includes: Acquire an original gesture image of the user; Performing image enhancement on the original gesture image to obtain an image-enhanced gesture image; Performing a filtering operation on the image-enhanced gesture image to obtain a filtered gesture image; the filtering operation includes any one of mean filtering, median filtering, and Gaussian filtering; A normalization operation is performed on the filtered gesture image to obtain the gesture image.
[0012] According to a video surveillance pan-tilt control method provided by the present invention, the method further includes determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image, and then further including: When the pan / tilt positioning gesture in the gesture recognition result is a hand-crossing gesture, user prompt information is displayed; the user prompt information is to align the center point of the camera with the hand-crossing point.
[0013] The present invention also provides a video surveillance PTZ control system, comprising the following units: An acquisition unit, configured to acquire a user's gesture image; An input unit, configured to input the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; a determination unit, configured to determine a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result including a start control gesture, an end control gesture, and a pan / tilt positioning gesture; the start control gesture is used to activate a pan / tilt control subsystem, the pan / tilt positioning gesture is used to control pan / tilt rotation of the pan / tilt control subsystem, and the end control gesture is used to stop pan / tilt rotation; A generating unit is used to generate a control instruction corresponding to the gesture recognition result.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-described video surveillance PTZ control methods is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the video surveillance pan-tilt control method as described above is implemented.
[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned video surveillance pan-tilt control methods.
[0017] The video surveillance pan-tilt control method, system, and electronic device provided by the present invention acquire a user's gesture image, input the gesture image into a depth estimation model, and obtain gesture depth information output by the depth estimation model. Then, based on the gesture depth information and the gesture image, a gesture recognition result corresponding to the gesture image is determined. The gesture recognition result includes a start control gesture, an end control gesture, and a pan-tilt positioning gesture. The start control gesture is used to activate the pan-tilt control subsystem, the pan-tilt positioning gesture is used to control the pan-tilt rotation of the pan-tilt control subsystem, and the end control gesture is used to stop the pan-tilt rotation. Finally, a control instruction corresponding to the gesture recognition result is generated. Compared with the inconvenience and complexity of existing pan-tilt control methods, gesture recognition based on gesture images using a gesture recognition model allows operators to precisely control the pan-tilt rotation remotely with simple gestures, improving the convenience and flexibility of pan-tilt control. Furthermore, gesture recognition is performed based on gesture depth information and gesture images. The gesture depth information provides the position and shape of the hand in three-dimensional space, thereby utilizing the complementarity of gesture depth information and gesture images to improve the accuracy of gesture recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is one of the flow charts of the video surveillance PTZ control method provided by the present invention.
[0020] Figure 2 This is the second flow chart of the video surveillance PTZ control method provided by the present invention.
[0021] Figure 3 This is the third flow chart of the video surveillance PTZ control method provided by the present invention.
[0022] Figure 4 This is the fourth flow chart of the video surveillance PTZ control method provided by the present invention.
[0023] Figure 5 It is a structural diagram of the video monitoring pan-tilt control system provided by the present invention.
[0024] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0026] The terms "first," "second," and the like in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type.
[0027] Figure 1 This is one of the flow charts of the video surveillance PTZ control method provided by the present invention. Figure 2 This is the second flow chart of the video surveillance PTZ control method provided by the present invention. Figure 3 This is the third flow chart of the video surveillance PTZ control method provided by the present invention, as shown in FIG. Figure 1 、 Figure 2 and Figure 3 As shown, the method includes step 110 , step 120 , step 130 and step 140 .
[0028] Step 110: Acquire the user's gesture image.
[0029] Specifically, the video surveillance PTZ control method in the embodiment of the present invention is applied to a video surveillance PTZ control system. The video surveillance PTZ control system may include an operation interface, a signal processor, a device controller, and a PTZ device.
[0030] After the video surveillance PTZ control system is started, the camera parameters of the gesture acquisition subsystem are initialized, including adjusting the shooting angle, resolution, and frame rate. The gesture recognition subsystem's gesture recognition model is then loaded to ensure that it is operational. The models in the gesture recognition subsystem include a depth estimation model and a gesture recognition model.
[0031] The PTZ control subsystem enters standby mode, waiting to receive a start control command; the feedback and display subsystem initializes the display interface and displays a system standby prompt message.
[0032] The user makes gestures within the camera shooting range of the gesture acquisition subsystem, and the camera collects the user's gesture image in real time and transmits it to the gesture recognition subsystem.
[0033] The gesture acquisition subsystem uses a high-resolution camera to clearly capture gestures, providing a rich data foundation for subsequent accurate gesture recognition. The camera is installed within the operator's field of view to ensure clear capture of various gestures.
[0034] Step 120: Input the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model.
[0035] Specifically, after the gesture image is acquired, the gesture image may be input into a depth estimation model to obtain gesture depth information output by the depth estimation model.
[0036] It should be noted that before inputting the gesture image into the depth estimation model and obtaining the gesture depth information output by the depth estimation model, the gesture image may be preprocessed, including but not limited to image enhancement, normalization, sharpening, denoising, etc. Normalization refers to mapping the data of each dimension of the data vector to the interval between (0, 1) or (-1, 1), or mapping a norm of the data vector to 1.
[0037] Sharpening here refers to compensating for the contours of gesture images, enhancing their edges and grayscale transitions, and making them clearer. This can be categorized into spatial domain processing and frequency domain processing. By highlighting the edges and contours of objects in gesture images, or the features of certain linear target elements, the contrast between the edges of objects and surrounding pixels is increased. Denoising is the process of reducing noise in digital images. Generally, the digitization and transmission of images are often affected by interference from imaging devices and external environmental noise. The received image information generally contains noise, which can be a significant source of image interference. Denoising gesture images can remove this noise, further improving the authenticity and accuracy of the resulting gesture images.
[0038] The above-mentioned pre-processing operation on the gesture image can scale the gesture image to an appropriate size and effectively improve the clarity of the gesture image itself, facilitating subsequent processing of the gesture image. Here, the pre-processing operation on the gesture image can be implemented by the gesture recognition subsystem.
[0039] Here, the depth estimation model may be a monocular depth estimation model, a binocular depth estimation model, or a video depth estimation model, etc., which is not specifically limited in the embodiment of the present invention.
[0040] Gesture depth information refers to depth information related to a gesture, which can be used to describe the position, shape, and movement of a gesture in three-dimensional space. Gesture depth information may include a depth map, which is a two-dimensional image in which the value of each pixel represents the distance from the pixel to the camera.
[0041] Step 130: Determine a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result includes a start control gesture, an end control gesture, and a pan / tilt positioning gesture; the start control gesture is used to activate a pan / tilt control subsystem, the pan / tilt positioning gesture is used to control pan / tilt rotation of the pan / tilt control subsystem, and the end control gesture is used to stop pan / tilt rotation; Step 140: Generate a control instruction corresponding to the gesture recognition result.
[0042] Specifically, after obtaining the gesture depth information and gesture image, the gesture recognition result corresponding to the gesture image can be determined based on the gesture depth information and gesture image. Here, determining the gesture recognition result corresponding to the gesture image based on the gesture depth information and gesture image can be achieved through a gesture recognition model. The gesture recognition results include a start control gesture, an end control gesture, and a gimbal positioning gesture. The start control gesture is used to activate the gimbal control subsystem, the gimbal positioning gesture is used to control the gimbal rotation of the gimbal control subsystem, and the end control gesture is used to stop the gimbal rotation.
[0043] In one embodiment, if the gesture recognition result is a start control gesture, the gimbal control subsystem is activated and enters an operational state. Subsequently, when a gimbal positioning gesture is recognized, the gimbal control subsystem generates a control instruction based on the gesture information and transmits the instruction to the gimbal device via the corresponding communication protocol, driving the gimbal motor to achieve gimbal positioning. If the gesture recognition result is an end control gesture, the gimbal control subsystem terminates the current operation, the gimbal maintains its current position, and the system returns to a standby state.
[0044] During the training phase of the gesture recognition model, a large amount of sample data covering three types of gestures—start control, end control, and gimbal positioning—is collected. This sample data includes gesture samples from different operators, lighting conditions, and background environments. This improves the model's generalization and adaptability to a variety of complex scenarios. When an operator makes a gesture in front of the camera, the gesture acquisition subsystem captures the gesture image in real time and transmits it to the gesture recognition subsystem. The gesture recognition subsystem first preprocesses the image, including image enhancement, noise reduction, and normalization, to improve the clarity and stability of the gesture image, facilitating subsequent feature extraction and recognition. The preprocessed gesture image is input into the trained gesture recognition model, which matches the gesture image features with the labeled gesture features in the training data and quickly and accurately outputs the recognition result, indicating the category corresponding to the gesture (start control, end control, or gimbal positioning).
[0045] In one embodiment, during PTZ control, the feedback and display subsystem acquires real-time PTZ rotation status information (e.g., rotation direction and angle) and visually presents it to the operator. Furthermore, at various stages of system operation, the feedback and display subsystem provides the operator with prompts based on actual conditions, such as system standby or gesture recognition failure.
[0046] The method provided by an embodiment of the present invention obtains a user's gesture image, then inputs the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model. Then, based on the gesture depth information and the gesture image, a gesture recognition result corresponding to the gesture image is determined. The gesture recognition result includes a start control gesture, an end control gesture, and a gimbal positioning gesture. The start control gesture is used to activate the gimbal control subsystem, the gimbal positioning gesture is used to control the gimbal rotation of the gimbal control subsystem, and the end control gesture is used to stop the gimbal rotation. Finally, a control instruction corresponding to the gesture recognition result is generated. Compared to the inconvenience and complexity of existing gimbal control methods, gesture recognition based on gesture images using a gesture recognition model allows operators to precisely control the gimbal rotation remotely with simple gestures, improving the convenience and flexibility of gimbal control. Furthermore, gesture recognition is performed based on gesture depth information and gesture images. The gesture depth information provides the position and shape of the hand in three-dimensional space, thereby leveraging the complementarity of gesture depth information and gesture images to improve gesture recognition accuracy.
[0047] Based on the above embodiment, step 130 includes: Step 131: inputting the gesture depth information and the gesture image into a gesture recognition model to obtain a gesture recognition result corresponding to the gesture image output by the gesture recognition model; The gesture recognition model includes a three-dimensional key point sequence extraction module, a gesture feature extraction module, a fusion module and a classifier; The three-dimensional key point sequence extraction module is used to extract the time sequence features corresponding to the three-dimensional key point sequence in the gesture depth information; The gesture feature extraction module is used to extract spatial features in the gesture image; The fusion module is used to determine a fusion feature based on the temporal feature and the spatial feature; The classifier is used to determine a gesture recognition result corresponding to the gesture image based on the fusion feature.
[0048] Specifically, the gesture depth information and the gesture image are input into a gesture recognition model, and a gesture recognition result corresponding to the gesture image output by the gesture recognition model is obtained.
[0049] Among them, the gesture recognition model includes a three-dimensional key point sequence extraction module, a gesture feature extraction module, a fusion module and a classifier.
[0050] The 3D key point sequence extraction module is used to extract temporal features corresponding to the 3D key point sequence in the gesture depth information. This module can be a bidirectional LSTM (Long Short-Term Memory) model or a Transformer model, though this is not specifically limited in this embodiment of the present invention.
[0051] The gesture feature extraction module is used to extract spatial features from gesture images. Here, the gesture feature extraction module can be a cascaded multi-layer convolutional neural network (CNN), a deep neural network (DNN), or a combination of CNN and DNN, etc., which is not specifically limited in the embodiments of the present invention.
[0052] The fusion module is used to determine the fusion features based on the temporal features and the spatial features. Here, the temporal features and the spatial features are fused to obtain the fusion features. The fusion features can be obtained by splicing the temporal features and the spatial features, or by weighting the temporal features and the spatial features using the attention mechanism and then splicing them. The embodiments of the present invention do not make specific limitations on this.
[0053] Here, the classifier is used to determine the gesture recognition result corresponding to the gesture image based on the fused features.
[0054] The method provided by an embodiment of the present invention inputs gesture depth information and a gesture image into a gesture recognition model to obtain a gesture recognition result corresponding to the gesture image output by the gesture recognition model. The gesture recognition model includes a three-dimensional key point sequence extraction module, a gesture feature extraction module, a fusion module, and a classifier. The three-dimensional key point sequence extraction module is used to extract temporal features corresponding to the three-dimensional key point sequence in the gesture depth information; the gesture feature extraction module is used to extract spatial features in the gesture image; the fusion module is used to determine fusion features based on temporal features and spatial features; and the classifier is used to determine the gesture recognition result corresponding to the gesture image based on the fusion features. By extracting the temporal features corresponding to the three-dimensional key point sequence, this method can more accurately capture the dynamic changes of gestures and support complex gesture recognition, such as hand rotation and bending. Moreover, by combining temporal features and spatial features, it can handle dynamic gesture recognition tasks and support the recognition of continuous gesture movements.
[0055] Based on the above embodiment, the gesture recognition result includes the coordinates of the finger key points and the gesture classification result; Step 131 includes: Step 1311 : Input the gesture depth information and the gesture image into a gesture recognition model to obtain the gesture key point coordinates and the gesture classification result output by the gesture recognition model.
[0056] Specifically, gesture recognition results include finger keypoint coordinates and gesture classification results. Accordingly, gesture depth information and gesture images are input into the gesture recognition model, which then outputs gesture keypoint coordinates and gesture classification results. Gesture keypoint coordinates refer to the positional information of key points describing a gesture (such as the hand and fingers) in three-dimensional space. These keypoints, including key locations like the wrist and finger joints, accurately describe the shape and posture of a gesture. Gesture classification results include start and end control gestures, as well as gimbal positioning gestures.
[0057] Based on the above embodiment, step 1311 further includes: Step 1311 - 1 : Determine the direction and angle of each finger based on the coordinates of the finger key points.
[0058] Specifically, after obtaining the coordinates of the finger key points, the direction and angle pointed by each finger can be determined based on the coordinates of the finger key points.
[0059] In a preferred embodiment, when the gesture recognition subsystem recognizes a start control gesture, the pan / tilt control subsystem is activated and enters an operational state, allowing the system to receive and execute subsequent pan / tilt positioning gestures. Upon recognizing an end control gesture, the pan / tilt control subsystem immediately ceases all current operations, maintaining the pan / tilt's current position, and entering a standby state, awaiting the next start control gesture. For the pan / tilt positioning function, specific hand movements are defined as positioning gestures. For example, an operator extends their index finger and points in a certain direction, which corresponds to the desired pan / tilt rotation direction. For example, pointing the index finger upwards rotates the pan / tilt upwards; pointing the index finger downwards rotates the pan / tilt downwards; and pointing the index finger left or right rotates the pan / tilt accordingly. Precise pan / tilt positioning can be achieved through specific hand gestures, such as crossing one's hands to indicate that the camera center point should be aligned with the intersection of the two hands. Specifically, if the pan / tilt positioning gesture in the gesture recognition result is a crossed hands gesture, a user prompt is displayed, in which the user prompt is to align the camera center point with the intersection of the two hands.
[0060] When a pan / tilt positioning gesture is recognized, the pan / tilt control subsystem generates corresponding control instructions based on the direction and angle information represented by the gesture. These control instructions are sent to the pan / tilt device, driving the pan / tilt motor to achieve precise horizontal and vertical rotation of the pan / tilt, thereby adjusting the camera's shooting angle to the operator's desired position.
[0061] During the PTZ's positioning operation, the feedback and display subsystem obtains real-time information about the PTZ's rotation status, including rotation direction and angle. This information is presented to the operator in a visual manner, such as by displaying a real-time rotation animation of the PTZ on the user interface, allowing the operator to intuitively understand the PTZ's operation. In the system's standby mode, the feedback and display subsystem can display prompts to guide the operator in making the correct start control gesture. During operation, if the operator's gesture cannot be correctly recognized, the system can also provide corresponding prompts through this subsystem to help the operator adjust the gesture.
[0062] The method provided by the embodiment of the present invention has (1) a significant improvement in convenience: users do not need to carry a remote control or use a computer or other device to perform complex operations. They can remotely control the camera pan / tilt through simple gestures, which greatly improves the convenience of operation. The method is suitable for use in elderly care scenarios, making it convenient for the elderly to control the camera angle. In addition, it is also suitable for use in factories, kitchens, and other scenarios where people who wear gloves to work cannot operate mobile phone software.
[0063] (2) Greatly enhanced flexibility: Freed from the limitations of cables and equipment locations in traditional control methods, operators can move freely and perform gesture operations as long as they are within the effective range of the gesture acquisition subsystem, which increases the flexibility of pan-tilt control and can adapt to various complex working environments and operational requirements.
[0064] (3) Improved control efficiency: Through real-time gesture recognition and fast control command transmission, the PTZ can achieve rapid response control, greatly shortening the time interval from the operator issuing a command to the PTZ executing an action, meeting the needs of application scenarios with high real-time requirements, such as traffic violation law enforcement.
[0065] Based on the above embodiment, step 1311-1 includes: Step 210, determining the direction pointed by each finger based on the fingertip coordinates and the finger base coordinates in the finger key point coordinates; Step 220, determining a first vector based on the middle finger joint coordinates and the fingertip coordinates in the finger key point coordinates, and determining a second vector based on the middle finger joint coordinates and the finger base coordinates; Step 230: Determine the angle pointed by each finger based on the angle between the first vector and the second vector.
[0066] Specifically, the direction of each finger can be determined based on the fingertip coordinates and finger root coordinates in the finger key point coordinates. Specifically, the vector from the fingertip coordinates to the finger root coordinates is calculated, and the direction of this vector is the direction of the finger.
[0067] Then, based on the coordinates of the middle finger joint and the fingertip in the finger key point coordinates, the first vector is determined, and based on the coordinates of the middle finger joint and the finger root coordinates, the second vector is determined. =fingertip coordinates-middle finger joint coordinates, = finger root coordinate - middle finger joint coordinate.
[0068] Finally, based on the angle between the first vector and the second vector, the angle pointed by each finger is determined. Specifically, the cosine value between the first vector and the second vector can be determined based on the first vector and the second vector, and the angle corresponding to the cosine value can be calculated using the inverse cosine function. This angle is the angle pointed by each finger.
[0069] The method provided by an embodiment of the present invention determines the direction pointed by each finger based on the fingertip coordinates and finger root coordinates in the finger key point coordinates, determines a first vector based on the middle finger joint coordinates and fingertip coordinates in the finger key point coordinates, and determines a second vector based on the middle finger joint coordinates and finger root coordinates. Finally, based on the angle between the first vector and the second vector, the angle pointed by each finger is determined. By determining the direction and angle of the finger based on the finger key point coordinates, this method can significantly improve the accuracy and robustness of gesture recognition, support complex gesture recognition and dynamic gesture recognition, enhance the naturalness of interaction, and is applicable to a variety of application scenarios. This method combines deep learning and mathematical formulas, can efficiently extract and analyze gesture features, and provides strong technical support for gesture recognition and human-computer interaction.
[0070] Based on the above embodiment, the training step of the gesture recognition model includes: Step 310: Acquire an initial gesture recognition model; the initial gesture recognition model includes an initial three-dimensional key point sequence extraction module, an initial gesture feature extraction module, an initial fusion module, a regression branch, and a classification branch; Step 320: Obtain a sample gesture image and a labeled gesture classification result of the sample gesture image; the labeled gesture classification result includes the coordinates of the labeled finger key points and the labeled gesture classification result; Step 330, determining sample gesture depth information of the sample gesture image; Step 340: extracting sample time sequence features corresponding to the 3D key point sequence in the sample gesture depth information based on the initial 3D key point sequence extraction module; Step 350: extracting sample spatial features from the sample gesture image based on the gesture feature extraction module; Step 360: performing feature fusion on the temporal features and the spatial features based on the fusion module to obtain a sample fusion feature; Step 370: Predicting the coordinates of the finger key points based on the sample fusion features of the sample, and obtaining the predicted coordinates of the finger key points; Step 380: performing gesture classification prediction on the sample fusion features based on the classification branch to obtain a predicted gesture classification result; Step 390: Determine a target loss based on the predicted gesture classification result and the labeled gesture classification result, as well as the predicted finger key point coordinates and the labeled finger key point coordinates, and perform parameter iteration on the initial gesture recognition model based on the target loss to obtain the gesture recognition model.
[0071] Specifically, first, an initial gesture recognition model may be obtained. Here, the parameters of the initial gesture recognition model may be pre-set or randomly generated, which is not specifically limited in the embodiment of the present invention.
[0072] Among them, the initial gesture recognition model includes an initial three-dimensional key point sequence extraction module, an initial gesture feature extraction module, an initial fusion module, a regression branch and a classification branch.
[0073] Then, a sample gesture image and a labeled gesture classification result of the sample gesture image are obtained, wherein the labeled gesture classification result includes the labeled finger key point coordinates and the labeled gesture classification result.
[0074] Determine the depth information of the sample gesture of the sample gesture image, which can be achieved by a depth estimation model.
[0075] Furthermore, based on the initial 3D key point sequence extraction module, sample temporal features corresponding to the 3D key point sequence in the sample gesture depth information are extracted. The initial 3D key point sequence extraction module can be a bidirectional LSTM model or a Transformer model, which is not specifically limited in this embodiment of the present invention.
[0076] Then, based on the gesture feature extraction module, the sample spatial features in the sample gesture image are extracted. The initial gesture feature extraction module can be a multi-layer convolutional neural network with a cascade structure, a deep neural network, or a combination of CNN and DNN, etc. The embodiment of the present invention does not specifically limit this.
[0077] Based on the fusion module, the temporal features and spatial features are fused to obtain sample fusion features, and then the finger key point coordinates are predicted based on the regression branch to obtain the predicted finger key point coordinates.
[0078] Based on the classification branch, the sample fusion features are used to predict the gesture classification and obtain the predicted gesture classification result.
[0079] Finally, after obtaining the predicted gesture classification result and the predicted finger key point coordinates, the classification loss can be determined based on the difference between the predicted gesture classification result and the labeled gesture classification result, and the regression loss can be determined based on the difference between the predicted finger key point coordinates and the labeled finger key point coordinates. Based on the regression loss and the classification loss, the target loss is determined. Based on the target loss, the parameters of the initial gesture recognition model are iterated, and the initial gesture recognition model after the parameter iteration is completed is used as the gesture recognition model.
[0080] It can be understood that the smaller the difference between the predicted finger key point coordinates and the labeled finger key point coordinates, the smaller the regression loss; the larger the difference between the predicted finger key point coordinates and the labeled finger key point coordinates, the larger the regression loss.
[0081] It can be understood that the smaller the difference between the predicted gesture classification result and the labeled gesture classification result, the smaller the classification loss; the larger the difference between the predicted gesture classification result and the labeled gesture classification result, the larger the classification loss.
[0082] Based on the above embodiment, step 110 includes: Step 111, obtaining the original gesture image of the user; Step 112: performing image enhancement on the original gesture image to obtain an image-enhanced gesture image; Step 113: performing a filtering operation on the image-enhanced gesture image to obtain a filtered gesture image; the filtering operation may include any one of mean filtering, median filtering, and Gaussian filtering; Step 114 : performing a normalization operation on the filtered gesture image to obtain the gesture image.
[0083] Specifically, first, an original gesture image of the user may be obtained, and image enhancement may be performed on the original gesture image to obtain an enhanced gesture image, wherein the image enhancement includes at least one of histogram equalization, gamma correction, and edge enhancement.
[0084] Then, a filtering operation is performed on the image-enhanced gesture image to obtain a filtered gesture image; the filtering operation includes any one of mean filtering, median filtering and Gaussian filtering.
[0085] Finally, the filtered gesture image is normalized to obtain a gesture image. During the application process, normalizing the input image can ensure that the image data is consistent with the data format and distribution during model training, thereby ensuring the performance of the model.
[0086] Based on any of the above embodiments, Figure 4 This is a fourth flow chart of the video surveillance PTZ control method provided by the present invention, as shown in FIG. Figure 4 As shown, the method includes: First, start the video surveillance PTZ control system and initialize each subsystem. For example, the gesture recognition subsystem loads the gesture recognition model, the PTZ control subsystem enters the standby state, the gesture acquisition subsystem initializes the camera parameters, and the feedback and display subsystem displays the standby prompt.
[0087] After the gesture acquisition subsystem initializes the camera parameters, the user makes a gesture. The gesture acquisition subsystem captures gesture images in real time and transmits the images to the gesture recognition subsystem. The gesture recognition subsystem preprocesses the images and inputs them into the depth estimation model and gesture recognition model to output the gesture recognition results.
[0088] Among them, the gesture recognition results include start control gesture, end control gesture and gimbal positioning gesture. In addition, there are invalid gestures, which prompt that the gesture is invalid and the feedback and display subsystem displays a standby prompt.
[0089] Here, when the gesture recognition result is a start control gesture, the gimbal control subsystem is activated and enters an operational state; when the gesture recognition result is an end control gesture, the gimbal operation is stopped, the current position of the gimbal is maintained, and the system returns to the standby state; when the gesture recognition result is a gimbal positioning gesture, a gimbal control instruction is generated, the control instruction is sent to the gimbal device, the gimbal motor rotates, the feedback and display subsystem obtains the gimbal status, and displays the gimbal rotation status.
[0090] The video surveillance pan-tilt control system provided by the present invention is described below. The video surveillance pan-tilt control system described below and the video surveillance pan-tilt control method described above can be referenced to each other.
[0091] Based on any of the above embodiments, the present invention provides a video surveillance PTZ control system, Figure 5 This is a schematic diagram of the structure of the video monitoring PTZ control system provided by the present invention. Figure 5 As shown, the system includes: An acquisition unit 510 is configured to acquire a user's gesture image; An input unit 520, configured to input the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; a determination unit 530 configured to determine a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result including a start control gesture, an end control gesture, and a pan / tilt positioning gesture; the start control gesture is used to activate a pan / tilt control subsystem, the pan / tilt positioning gesture is used to control pan / tilt rotation of the pan / tilt control subsystem, and the end control gesture is used to stop pan / tilt rotation; The generating unit 540 is configured to generate a control instruction corresponding to the gesture recognition result.
[0092] The system provided by an embodiment of the present invention obtains a user's gesture image, then inputs the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model. Then, based on the gesture depth information and the gesture image, a gesture recognition result corresponding to the gesture image is determined. The gesture recognition result includes a start control gesture, an end control gesture, and a pan / tilt positioning gesture. The start control gesture is used to activate the pan / tilt control subsystem, the pan / tilt positioning gesture is used to control the pan / tilt rotation of the pan / tilt control subsystem, and the end control gesture is used to stop the pan / tilt rotation. Finally, a control instruction corresponding to the gesture recognition result is generated. Compared with the inconvenience and complexity of existing pan / tilt control methods, gesture recognition based on gesture images using a gesture recognition model allows operators to precisely control the pan / tilt rotation remotely with simple gestures, improving the convenience and flexibility of pan / tilt control. Furthermore, gesture recognition is performed based on gesture depth information and gesture images. The gesture depth information provides the position and shape of the hand in three-dimensional space, thereby utilizing the complementarity of gesture depth information and gesture images to improve the accuracy of gesture recognition.
[0093] Based on any of the above embodiments, the determining unit 530 specifically includes: a gesture recognition unit, configured to input the gesture depth information and the gesture image into a gesture recognition model, and obtain a gesture recognition result corresponding to the gesture image output by the gesture recognition model; The gesture recognition model includes a three-dimensional key point sequence extraction module, a gesture feature extraction module, a fusion module and a classifier; The three-dimensional key point sequence extraction module is used to extract the time sequence features corresponding to the three-dimensional key point sequence in the gesture depth information; The gesture feature extraction module is used to extract spatial features in the gesture image; The fusion module is used to determine a fusion feature based on the temporal feature and the spatial feature; The classifier is used to determine a gesture recognition result corresponding to the gesture image based on the fusion feature.
[0094] Based on any of the above embodiments, the gesture recognition result includes the coordinates of the finger key points and the gesture classification result; The gesture recognition unit is specifically used to: The gesture depth information and the gesture image are input into a gesture recognition model to obtain the gesture key point coordinates and the gesture classification result output by the gesture recognition model.
[0095] Based on any of the above embodiments, the method further includes a direction and angle determination unit, wherein the direction and angle determination unit specifically includes: The determination subunit is used to determine the direction and angle of each finger based on the coordinates of the finger key points.
[0096] Based on any of the foregoing embodiments, the determining subunit is specifically configured to: Determining the direction pointed by each finger based on the fingertip coordinates and the finger base coordinates in the finger key point coordinates; Determine a first vector based on the middle finger joint coordinates and the fingertip coordinates in the finger key point coordinates, and determine a second vector based on the middle finger joint coordinates and the finger base coordinates; The angle pointed by each finger is determined based on the angle between the first vector and the second vector.
[0097] Based on any of the above embodiments, the further comprising a training unit, wherein the training unit is specifically configured to: Acquire an initial gesture recognition model; the initial gesture recognition model includes an initial three-dimensional key point sequence extraction module, an initial gesture feature extraction module, an initial fusion module, a regression branch, and a classification branch; Acquire a sample gesture image and a labeled gesture classification result of the sample gesture image; the labeled gesture classification result includes the labeled finger key point coordinates and the labeled gesture classification result; determining sample gesture depth information of the sample gesture image; Extracting sample time sequence features corresponding to the three-dimensional key point sequence in the sample gesture depth information based on the initial three-dimensional key point sequence extraction module; Extracting sample spatial features from the sample gesture image based on the gesture feature extraction module; Performing feature fusion on the temporal features and the spatial features based on the fusion module to obtain a sample fusion feature; Predicting the coordinates of the finger key points on the sample fusion features based on the regression branch to obtain the predicted coordinates of the finger key points; Performing gesture classification prediction on the sample fusion features based on the classification branch to obtain a predicted gesture classification result; Based on the predicted gesture classification result and the labeled gesture classification result, as well as the predicted finger key point coordinates and the labeled finger key point coordinates, a target loss is determined, and parameters of the initial gesture recognition model are iterated based on the target loss to obtain the gesture recognition model.
[0098] Based on any of the above embodiments, the acquiring unit 510 is specifically configured to: Acquire an original gesture image of the user; Performing image enhancement on the original gesture image to obtain an image-enhanced gesture image; Performing a filtering operation on the image-enhanced gesture image to obtain a filtered gesture image; the filtering operation includes any one of mean filtering, median filtering, and Gaussian filtering; A normalization operation is performed on the filtered gesture image to obtain the gesture image.
[0099] Based on any of the above embodiments, the further embodiment includes a prompting unit, wherein the prompting unit is specifically configured to: When the pan / tilt positioning gesture in the gesture recognition result is a hand-crossing gesture, user prompt information is displayed; the user prompt information is to align the center point of the camera with the hand-crossing point.
[0100] Figure 6 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communications bus 640. The processor 610, the communications interface 620, and the memory 630 communicate with each other via the communications bus 640. The processor 610 may invoke logic instructions in the memory 630 to execute a video surveillance pan-tilt control method, which includes: obtaining a user's gesture image; inputting the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result includes a start control gesture, an end control gesture, and a pan-tilt positioning gesture; the start control gesture is used to activate the pan-tilt control subsystem, the pan-tilt positioning gesture is used to control pan-tilt rotation of the pan-tilt control subsystem, and the end control gesture is used to stop pan-tilt rotation; and generating a control instruction corresponding to the gesture recognition result.
[0101] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0102] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the video surveillance pan-tilt control method provided by the above methods, the method including: obtaining a user's gesture image; inputting the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result includes a start control gesture, an end control gesture and a pan-tilt positioning gesture; the start control gesture is used to activate the pan-tilt control subsystem, the pan-tilt positioning gesture is used to control the pan-tilt rotation of the pan-tilt control subsystem, and the end control gesture is used to stop the pan-tilt rotation; and generating a control instruction corresponding to the gesture recognition result.
[0103] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the video surveillance pan-tilt control method provided by the above-mentioned methods, the method comprising: obtaining a user's gesture image; inputting the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result comprises a start control gesture, an end control gesture and a pan-tilt positioning gesture; the start control gesture is used to activate the pan-tilt control subsystem, the pan-tilt positioning gesture is used to control the pan-tilt rotation of the pan-tilt control subsystem, and the end control gesture is used to stop the pan-tilt rotation; and generating a control instruction corresponding to the gesture recognition result.
[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0105] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A video surveillance PTZ control method, characterized in that: include: Get the user's gesture image; Inputting the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; Determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result includes a start control gesture, an end control gesture, and a gimbal positioning gesture; the start control gesture is used to activate a gimbal control subsystem, the gimbal positioning gesture is used to control the gimbal rotation of the gimbal control subsystem, and the end control gesture is used to stop the gimbal rotation; Generate a control instruction corresponding to the gesture recognition result.
2. The video surveillance PTZ control method according to claim 1, wherein: The determining, based on the gesture depth information and the gesture image, a gesture recognition result corresponding to the gesture image includes: Inputting the gesture depth information and the gesture image into a gesture recognition model to obtain a gesture recognition result corresponding to the gesture image output by the gesture recognition model; The gesture recognition model includes a three-dimensional key point sequence extraction module, a gesture feature extraction module, a fusion module and a classifier; The three-dimensional key point sequence extraction module is used to extract the time sequence features corresponding to the three-dimensional key point sequence in the gesture depth information; The gesture feature extraction module is used to extract spatial features in the gesture image; The fusion module is used to determine a fusion feature based on the temporal feature and the spatial feature; The classifier is used to determine a gesture recognition result corresponding to the gesture image based on the fusion feature.
3. The video surveillance PTZ control method according to claim 2, wherein: The gesture recognition result includes the coordinates of the finger key points and the gesture classification result; Inputting the gesture depth information and the gesture image into a gesture recognition model to obtain a gesture recognition result corresponding to the gesture image output by the gesture recognition model includes: The gesture depth information and the gesture image are input into a gesture recognition model to obtain the gesture key point coordinates and the gesture classification result output by the gesture recognition model.
4. The video surveillance PTZ control method according to claim 3, wherein: The step of inputting the gesture depth information and the gesture image into a gesture recognition model to obtain the gesture key point coordinates and the gesture classification result output by the gesture recognition model further includes: Based on the coordinates of the finger key points, the direction and angle of each finger are determined.
5. The video surveillance PTZ control method according to claim 4, characterized in that: The determining the direction and angle of each finger based on the coordinates of the finger key points includes: Determining the direction pointed by each finger based on the fingertip coordinates and the finger base coordinates in the finger key point coordinates; Determine a first vector based on the middle finger joint coordinates and the fingertip coordinates in the finger key point coordinates, and determine a second vector based on the middle finger joint coordinates and the finger base coordinates; The angle pointed by each finger is determined based on the angle between the first vector and the second vector.
6. The video surveillance PTZ control method according to claim 2, wherein: The training steps of the gesture recognition model include: Acquire an initial gesture recognition model; the initial gesture recognition model includes an initial three-dimensional key point sequence extraction module, an initial gesture feature extraction module, an initial fusion module, a regression branch, and a classification branch; Acquire a sample gesture image and a labeled gesture classification result of the sample gesture image; the labeled gesture classification result includes the labeled finger key point coordinates and the labeled gesture classification result; determining sample gesture depth information of the sample gesture image; Extracting sample time sequence features corresponding to the three-dimensional key point sequence in the sample gesture depth information based on the initial three-dimensional key point sequence extraction module; Extracting sample spatial features from the sample gesture image based on the gesture feature extraction module; Performing feature fusion on the temporal features and the spatial features based on the fusion module to obtain a sample fusion feature; Predicting the coordinates of the finger key points on the sample fusion features based on the regression branch to obtain the predicted coordinates of the finger key points; Performing gesture classification prediction on the sample fusion features based on the classification branch to obtain a predicted gesture classification result; Based on the predicted gesture classification result and the labeled gesture classification result, as well as the predicted finger key point coordinates and the labeled finger key point coordinates, a target loss is determined, and parameters of the initial gesture recognition model are iterated based on the target loss to obtain the gesture recognition model.
7. The video surveillance PTZ control method according to any one of claims 1 to 6, characterized in that: The acquiring of the user's gesture image includes: Acquire an original gesture image of the user; Performing image enhancement on the original gesture image to obtain an image-enhanced gesture image; Performing a filtering operation on the image-enhanced gesture image to obtain a filtered gesture image; the filtering operation includes any one of mean filtering, median filtering, and Gaussian filtering; A normalization operation is performed on the filtered gesture image to obtain the gesture image.
8. The video surveillance PTZ control method according to any one of claims 1 to 6, characterized in that: The step of determining a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image further includes: When the pan / tilt positioning gesture in the gesture recognition result is a hand-crossing gesture, user prompt information is displayed; the user prompt information is to align the center point of the camera with the hand-crossing point.
9. A video surveillance PTZ control system, characterized in that: include: An acquisition unit, configured to acquire a user's gesture image; An input unit, configured to input the gesture image into a depth estimation model to obtain gesture depth information output by the depth estimation model; a determination unit, configured to determine a gesture recognition result corresponding to the gesture image based on the gesture depth information and the gesture image; the gesture recognition result including a start control gesture, an end control gesture, and a pan / tilt positioning gesture; the start control gesture is used to activate a pan / tilt control subsystem, the pan / tilt positioning gesture is used to control pan / tilt rotation of the pan / tilt control subsystem, and the end control gesture is used to stop pan / tilt rotation; A generating unit is used to generate a control instruction corresponding to the gesture recognition result.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the video surveillance pan-tilt control method according to any one of claims 1 to 8 is implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video surveillance pan-tilt control method according to any one of claims 1 to 8 is implemented.