Shooting target identification method and device, computer equipment and storage medium
By determining the consistency of the panoramic camera's pointing and face orientation, establishing a coordinate system and determining the search area, and placing the target in the center of the panoramic image, the problem of difficult target alignment during panoramic camera shooting is solved, and the accuracy of user experience and target recognition is improved.
Patent Information
- Application Number
- CN202510323795.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
AI Technical Summary
在使用全景相机拍摄时,用户难以将目标物置于图像中心,导致后续剪辑效果不佳。
By determining whether the direction of the panoramic camera is consistent with the direction of the photographer's face, using inertial sensor data and face detection algorithms, a coordinate system is established and the search area is determined, and the recognition target is placed in the center area of the panoramic image.
Improve user experience, ensuring that the focus of the photographer is at the center of the image, conforms to visual habits, and understands the scene and focus more accurately in scenes that are inconvenient to aligning the target.
Smart Images

Figure CN120302145A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of target recognition, and particularly to a method, apparatus, computer device, and storage medium for photographing target recognition. Background Art
[0002] With the continuous development of various photographing devices, users will use photographing devices to record various current contents in more and more scenarios. During the photographing process, the user will aim the camera at or close to the target object for photographing.
[0003] However, in some scenarios, it is inconvenient for the user to aim the photographing device at the target object. If the photographing device used is a panoramic camera, since the target object is not locked and cannot be placed at the center position, the subsequent editing effect will be affected during the editing process. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium for photographing target recognition.
[0005] In a first aspect, the present disclosure provides a method for photographing target recognition. The method includes:
[0006] Obtain a panoramic image obtained by a photographer through a panoramic camera;
[0007] Based on the panoramic image obtained by the panoramic camera and the position of the photographer in space, determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer;
[0008] In response to the pointing direction of the panoramic camera being consistent with the face orientation of the photographer, based on the pointing direction of the panoramic camera, determine a search area in the panoramic image;
[0009] Based on the search area, determine an identification target of the current panoramic image, and place the identification target in the central area of the panoramic image.
[0010] In one embodiment, the determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer based on the panoramic image obtained by the panoramic camera and the position of the photographer in space includes:
[0011] Determine the inertial sensor data currently corresponding to the panoramic image, determine the face of the photographer in the panoramic image, and establish a face coordinate system based on the face orientation and the gravity direction indicated by the inertial sensor data;
[0012] Project the panoramic camera into the face coordinate system, and determine a first pose of the panoramic camera in the face coordinate system relative to the face;
[0013] Determine the pointing direction of the panoramic camera in the first posture, and determine the first included angle between the face orientation and the pointing direction according to the planar projection of the pointing direction in the face coordinate system;
[0014] Based on the first included angle and a preset included angle threshold, determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer.
[0015] In one embodiment, the establishing a face coordinate system based on the face orientation and the gravity direction indicated by the inertial sensor data includes:
[0016] Determine a first coordinate axis according to the face orientation, determine a second coordinate axis according to the gravity direction, and determine a third coordinate axis by determining a straight line perpendicular to the first coordinate axis and the second coordinate axis;
[0017] Taking the face as the origin, establish a face coordinate system based on the first coordinate axis, the second coordinate axis, and the third coordinate axis.
[0018] In one embodiment, the determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer based on the panoramic image captured by the panoramic camera and the position of the photographer in space includes:
[0019] Determine the inertial sensor data currently corresponding to the panoramic image, and establish a camera coordinate system based on the gravity direction indicated by the inertial sensor data;
[0020] Project the panoramic image according to the camera coordinate system, determine the face orientation in the panoramic image, and calculate the Yaw pose angle according to the face orientation;
[0021] Based on the Yaw pose angle and a preset included angle threshold, determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer.
[0022] In one embodiment, the establishing a camera coordinate system based on the gravity direction indicated by the inertial sensor data includes:
[0023] Determine a first coordinate axis according to the gravity direction, determine a second coordinate axis according to the horizontal direction perpendicular to the gravity direction, and determine a third coordinate axis according to the gravity direction and the horizontal direction;
[0024] Taking the panoramic camera as the origin, establish a camera coordinate system according to the first coordinate axis, the second coordinate axis, and the third coordinate axis.
[0025] In one embodiment, determining a search area in a panoramic image based on the pointing direction of the panoramic camera includes:
[0026] Project the panoramic image to obtain a planar panoramic image;
[0027] Based on the pointing direction of the panoramic camera and the planar panoramic image, determine an initial pointing point in the planar panoramic image;
[0028] Based on the initial pointing point and a preset search range, determine the search area in the planar panoramic image.
[0029] In a second aspect, the present disclosure also provides a shooting target recognition device. The device includes:
[0030] A panoramic image acquisition module, configured to acquire a panoramic image captured by a panoramic camera by a photographer;
[0031] A face orientation detection module, configured to determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer based on the panoramic image captured by the panoramic camera and the position of the photographer in space;
[0032] A search area determination module, configured to, in response to the pointing direction of the panoramic camera being consistent with the face orientation of the photographer, determine a search area in the panoramic image based on the pointing direction of the panoramic camera;
[0033] A target recognition adjustment module, configured to determine a recognition target for the current panoramic image based on the search area and place the recognition target in the central area of the panoramic image.
[0034] In a third aspect, the present disclosure also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in any of the above method embodiments are implemented.
[0035] In a fourth aspect, the present disclosure also provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0036] In a fifth aspect, the present disclosure also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0037] In the above embodiments, by determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer, the intention of the photographer can be better understood. When the two are consistent, the search area is determined based on the camera pointing direction and the recognition target is placed in the central area of the panoramic image, so that the key content that the photographer focuses on can be in the center of the image, which conforms to the user's visual habits and expectations, improves the user experience when viewing the panoramic image, and the user can quickly find the content of interest without manual adjustment. In addition, in some scenarios where it is not convenient to directly align the panoramic camera with the target, combining the face orientation of the photographer can more accurately understand the scene and the focus of the photographer, so as to determine the recognition target. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart of a shooting target recognition method in an embodiment;
[0040] Figure 2 It is a schematic flowchart of step S104 in an embodiment;
[0041] Figure 3 It is a schematic flowchart of step S202 in an embodiment;
[0042] Figure 4 It is another schematic flowchart of step S104 in an embodiment;
[0043] Figure 5 It is a schematic flowchart of step S402 in an embodiment;
[0044] Figure 6 It is a schematic flowchart of step S106 in an embodiment;
[0045] Figure 7 It is a schematic block diagram of the structure of a shooting target recognition device in an embodiment;
[0046] Figure 8 It is a schematic internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0048] It should be noted that the terms "first", "second", etc. in the description and claims of this article and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or equipment.
[0049] In this article, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0050] In one embodiment, as Figure 1 shown, a method for identifying a shooting target is provided. In this embodiment, it is exemplified that the method is applied to a terminal, and includes the following steps:
[0051] S102, obtain a panoramic image obtained by the shooter through a panoramic camera.
[0052] S104, based on the panoramic image obtained by the panoramic camera and the position of the shooter in space, determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the shooter.
[0053] Among them, the position of the photographer in space can refer to the three-dimensional coordinates (x, y, z) where the photographer is located. The pointing direction of the panoramic camera can refer to the spatial direction that the panoramic camera lens is aimed at, usually described by Euler angles (Yaw, Pitch, Roll) or unit vectors. For example, with the gravity direction as a reference (assuming the gravity direction is the z-axis), combined with the rotation attitude of the camera, the pointing of the camera in the horizontal and vertical directions is determined. In practical applications, the built-in inertial measurement unit (IMU) of the camera can be used to measure acceleration and angular velocity, and then the attitude information of the camera can be calculated to clarify its pointing direction. The face orientation of the photographer can refer to the direction representing the front of the photographer's face, generally measured by an angle on the horizontal plane, called the Yaw angle.
[0054] Specifically, the panoramic camera continuously captures panoramic images. Advanced face detection algorithms (such as MTCNN, RetinaFace, etc.) are used to process the panoramic images to detect the face area of the photographer. Then, within the detected face area, with the help of face key-point detection techniques (such as face key-point detection in OpenCV, face mesh model in MediaPipe), the key feature points of the face, such as both eyes, the tip of the nose, the corners of the mouth, etc., are accurately located. According to the positional relationships of these feature points, the Yaw angle of the face is calculated through specific calculation methods (such as based on geometric relationships or machine learning models), thereby determining the face orientation of the photographer. The acceleration and angular velocity data are obtained through the built-in IMU of the panoramic camera. Integral operations are performed on these data, combined with the gravity direction (usually the gravity direction is used as the reference axis), to calculate the attitude information of the camera, and then the Euler angles (Yaw, Pitch, Roll) of the camera are obtained, where the Yaw angle represents the pointing of the panoramic camera in the horizontal direction. The SLAM technology can be used to determine the position of the photographer in space based on the movement trajectory of the photographer in the camera image and environmental feature points. The pointing direction of the panoramic camera (represented by the Yaw angle) and the face orientation of the photographer (also represented by the Yaw angle) are unified into the same coordinate system, and the included angle between them is calculated. According to the included angle, it is determined whether the pointing direction of the panoramic camera and the face orientation of the photographer are consistent.
[0055] S106, in response to the pointing direction of the panoramic camera being consistent with the face orientation of the photographer, based on the pointing direction of the panoramic camera, determine the search area in the panoramic image.
[0056] Among them, the search area can be a specific area preset in the panoramic image to find a specific target, which improves the target detection efficiency by narrowing the search range.
[0057] Specifically, the attitude data of the panoramic camera is obtained by using the IMU, the Yaw angle of the camera is calculated, and its pointing direction is determined. At the same time, from the images captured by the panoramic camera, faces are detected using face detection algorithms (such as MTCNN, RetinaFace), and then facial feature points are located through face key point detection (such as the MediaPipe face mesh model), the Yaw angle of the face is calculated to obtain the face orientation. A threshold angle (such as 10° or 15°) is set, and the Yaw angles of the camera and the face are compared. If the difference between the two is less than the threshold, it is determined that the pointing direction of the panoramic camera is consistent with the face orientation of the photographer. If it is detected that the camera pointing and the face orientation are consistent, the search area is determined based on the pointing direction of the camera.
[0058] S108, determining the recognition target of the current panoramic image based on the search area, and placing the recognition target in the central area of the panoramic image.
[0059] Among them, the recognition target is usually an object expected to be detected and recognized in the panoramic image, which can be a person, an object, a specific identifier, etc. Central area: The central part of the panoramic image, usually the area with higher visual attention in the image. For different application scenarios and image sizes, the definition and range of the central area may be different.
[0060] Specifically, within the determined search area, a target detection algorithm is used to identify the target. Commonly used target detection algorithms include deep learning-based algorithms such as the YOLO (You Only Look Once) series, Faster R-CNN, etc., as well as traditional target detection algorithms such as Haar feature cascade detection. If the recognition target is a face, dedicated face detection algorithms such as MTCNN (Multi-task Cascaded Convolutional Networks), RetinaFace, etc. can also be used. The image data within the search area is input into the selected target detection algorithm model, and the model will classify and locate each object in the image, outputting information such as the detected target category, position coordinates (such as the upper left and lower right coordinates of the rectangular box), etc. For example, using the YOLOv5 model, it will quickly scan within the search area, identify various targets that meet the definitions in the pre-trained model, and give the position and category information of each target. According to the target position information (such as the center coordinates of the rectangular box) output by the target detection algorithm and the coordinates of the central area of the panoramic image, the offset of the target relative to the central area is calculated. According to the calculated offset, the target is placed in the central area.
[0061] In some exemplary embodiments, for example, the central coordinates of the panoramic image are (X center , y center ), and the coordinate information of the recognition target is (xtarget , y target ), then the horizontal offset is: dx = X center - x target , and the vertical offset is: dy = y center - y target . If the panoramic image is processed in real time with software or hardware support, transformation operations such as image translation and rotation can be used. For example, the function in the OpenCV library is used to implement image translation, and the entire panoramic image is translated according to the offset to move the target to the central area. Specifically, by constructing a translation matrix and using the cv2.warpAffine function to transform the panoramic image.
[0062] In the above method for identifying the shooting target, by determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the shooter, the shooter's intention can be better understood. When the two are consistent, the search area is determined based on the camera pointing direction and the recognition target is placed in the central area of the panoramic image, so that the key content that the shooter focuses on can be in the center of the image, which conforms to the user's visual habits and expectations, and improves the user experience when viewing the panoramic image. The user can quickly find the content of interest without manual adjustment. In addition, in some scenarios where it is not convenient to directly use the panoramic camera to aim at the target, combining the face orientation of the shooter can more accurately understand the scene and the shooter's focus of attention, so as to determine the recognition target.
[0063] In one embodiment, as Figure 2 shown, determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the shooter based on the panoramic image obtained by the panoramic camera and the position of the shooter in space includes:
[0064] S202. Determine the inertial sensor data currently corresponding to the panoramic image, determine the face of the shooter in the panoramic image, and establish a face coordinate system based on the face orientation and the gravity direction indicated by the inertial sensor data.
[0065] Among them, inertial sensor data are usually the data collected by an inertial sensor (such as an inertial measurement unit IMU composed of an accelerometer, a gyroscope, etc.). The accelerometer can measure the acceleration of an object in three axes, which is used to determine the direction of gravity and the linear acceleration of the device; the gyroscope measures the angular velocity of the object around three axes, reflecting the rotational motion of the device. These data are used to describe the attitude and motion state of the panoramic camera. The direction of gravity usually refers to the direction of the earth's gravitational force. In inertial sensor data, the accelerometer can measure the gravitational acceleration vector to determine the direction of gravity. When establishing a coordinate system and analyzing the attitude of an object, the direction of gravity is often used as an important reference benchmark. The face coordinate system can be a coordinate system established with the face as the center, which is used to describe the direction and position information related to the face. The definition of its coordinate axes is usually related to factors such as the direction of gravity and the face orientation, so as to analyze the attitude and actions of the face more accurately.
[0066] Specifically, the inertial measurement unit (IMU) built into the panoramic camera can continuously collect accelerometer and gyroscope data. While shooting a panoramic image, the inertial sensor data collected by the IMU at the same moment is recorded. The inertial sensor data can be stored in the camera's memory in the form of digital signals or transmitted in real-time through a data interface to a connected processing device (such as a mobile phone or a computer). In the processing device, through the corresponding driver program and data reading function, these inertial sensor data are obtained and parsed to obtain the acceleration and angular velocity information at a specific time point, and this information is the inertial sensor data corresponding to the current panoramic image. The panoramic image is input into a pre-selected face detection algorithm model. The model analyzes each region in the image, extracts image features through a convolutional neural network (CNN), and determines whether each region contains a face. If a face is detected, the algorithm outputs the position information of the face, usually represented in the form of a rectangular box, including the upper left corner coordinates and the lower right corner coordinates of the rectangular box, so as to determine the position of the photographer's face in the panoramic image. The gravitational acceleration vector is extracted from the accelerometer data in the inertial sensor data. When the camera is stationary or the motion state is relatively stable, the gravitational acceleration dominates in the acceleration vector measured by the accelerometer. By processing the acceleration data of the three axes (assumed to be the x, y, and z axes) measured by the accelerometer, for example, removing the possible linear acceleration components (which can be achieved through some filtering algorithms), the gravitational acceleration vector is obtained. Within the detected face region, face key point detection techniques (such as MediaPipe's face mesh model, OpenCV's face key point detection) are used to locate the key feature points of the face, such as both eyes, the tip of the nose, the corners of the mouth, etc. According to the positional relationship of these feature points, the Yaw angle of the face is calculated through a specific calculation method (such as based on geometric relationships or machine learning models) to determine the orientation of the face in the horizontal direction. Taking the center of the face (such as the midpoint of the line connecting both eyes and the midpoint of the line connecting the tip of the nose) as the coordinate origin, a face coordinate system is established based on the face orientation and the gravitational direction indicated by the inertial sensor.
[0067] S204, project the panoramic camera into the face coordinate system, and determine the first pose of the panoramic camera relative to the face in the face coordinate system.
[0068] Among them, the first pose generally refers to the position and orientation information of the panoramic camera in the face coordinate system, which may include translation information (the position offsets of the camera relative to the face in the x, y, and z-axis directions) and rotation information (the rotation angles of the camera around the x, y, and z axes, usually represented by the Euler angles Yaw, Pitch, and Roll). Projection: In the sense of mathematics and computer graphics, the projection in the embodiments of the present disclosure refers to converting the position and pose information of the panoramic camera in a three-dimensional space into a specific three-dimensional space reference system, i.e., the face coordinate system, according to certain rules, so as to describe the position and orientation of the camera in this reference system.
[0069] Specifically, the directions of each coordinate axis in the face coordinate system can be determined, and the position coordinates of the panoramic camera in the global coordinate system can be determined. The position coordinates of the panoramic camera in the global coordinate system are converted into the position coordinates in the face coordinate system to determine the first pose of the panoramic camera relative to the face in the face coordinate system.
[0070] In some exemplary embodiments, for example, the coordinates of the origin of the face coordinate system in the global coordinate system are (x face , y face , z face ), and the position coordinates of the panoramic camera in the global coordinate system are (x cam , y cam , z cam ), then the translation vector of the panoramic camera in the face coordinate system is (x cam - x face , y cam - y face , z cam - z face ). The position offsets of the camera relative to the face in the three coordinate axis directions are calculated.
[0071] It is necessary to represent the Euler angles of the pose of the panoramic camera in the global coordinate system as and convert them to the face coordinate system. This involves the calculation and transformation of the rotation matrix. First, according to the conversion formula between the Euler angles and the rotation matrix, the rotation matrices around the x, y, and z axes in the global coordinate system are calculated respectively Then they are combined into the total rotation matrix in the global coordinate system. Next, the rotation matrix R face of the face coordinate system relative to the global coordinate system is calculated, and its inverse matrix can be used to convert the rotation information in the global coordinate system to the face coordinate system. Through matrix multiplication the rotation matrix of the panoramic camera in the face coordinate system is obtained, and then the Euler angles are extracted from this rotation matrix. These are the rotation angles of the panoramic camera around the x, y, and z axes in the face coordinate system.
[0072] Combining the results of the above translation transformation and rotation transformation, the first pose of the panoramic camera in the face coordinate system is obtained. The translation vector (x cam -x face , y cam -y face , z cam -z face ) describes the position of the camera relative to the face, and the Euler angles describe the rotation direction of the camera relative to the face. In this way, the first pose of the panoramic camera in the face coordinate system is completely determined.
[0073] S206. Determine the pointing direction of the panoramic camera in the first pose. According to the planar projection of the pointing direction in the face coordinate system, determine the first angle between the face orientation and the pointing direction.
[0074] Among them, the planar projection refers to the operation of projecting a vector in three-dimensional space (here it is the pointing direction vector of the panoramic camera) onto a two-dimensional plane. In this scenario, it is to project the pointing direction vector of the panoramic camera onto a certain plane in the face coordinate system (usually choose the horizontal plane containing the face orientation, which is convenient for calculating the angle with the face orientation). After projection, the three-dimensional vector is transformed into a two-dimensional vector, which simplifies the subsequent angle calculation process. The first angle refers to the angle between the face orientation vector and the planar projection vector of the pointing direction of the panoramic camera in the face coordinate system plane, and is used to measure the consistency degree between the face orientation and the camera pointing direction. By calculating this angle, it can be determined whether the camera is aligned with the face, providing a basis for subsequent operations such as shooting angle adjustment and target tracking. The pointing direction of the panoramic camera can represent the three-dimensional space direction aimed at by the panoramic camera lens. In the context of the current face coordinate system, it is determined based on the rotation information in the first pose. It can be understood as the direction vector along the optical axis of the lens starting from the optical center of the camera, and this vector has a specific coordinate representation in the face coordinate system.
[0075] Specifically, since the first pose of the panoramic camera relative to the face coordinate system has been determined, the pointing direction of the panoramic camera can be determined according to the first pose. Determine the planar projection of the pointing direction on the plane in the face coordinate system (usually the XY plane) to determine the first angle between the face orientation and the pointing direction.
[0076] In some exemplary embodiments, given the first pose of the panoramic camera in the face coordinate system, extract the rotation information, that is, the Euler angles (Yaw, Pitch, Roll). Assume that the Euler angles in the first pose are (θ y , θ p , θ r), respectively representing the rotation angles around the y-axis, p-axis, and r-axis. According to the conversion relationship between Euler angles and direction vectors, the pointing direction vector of the panoramic camera is constructed in the face coordinate system. In the Cartesian coordinate system, if the camera is taken as the origin, the x, y, and z components of the pointing direction vector can be calculated through trigonometric functions. For example, in a simple case (assuming that the initial directions of the camera coordinate system and the face coordinate system are the same), the z component of the pointing direction vector can be expressed as z = cos(θ p )cos(θ y ), the x component is |x = sin(θ y )cos(θ p ), and the y component is y = sin(θ p ), thus obtaining the pointing direction vector
[0077] Usually, the horizontal plane where the face is located (assumed to be the x-y plane) is selected for projection. This is because the face orientation is mainly considered in the horizontal direction, and such a choice facilitates subsequent calculation of the included angle.
[0078] For the pointing direction vector When projected onto the x-y plane, the z component of its projection vector becomes 0, and the x and y components remain unchanged, that is In the face coordinate system, assuming that the rotation angle of the face orientation around the z-axis on the horizontal plane is θ face , then the face orientation vector On the x-y plane can be expressed as (assuming that the modulus of the face orientation vector is 1).
[0079] Use the vector dot product formula To calculate the first included angle α. Among them, Is the dot product of two vectors, equal to xcos(θ face ) + ysin(θ face ), Substitute these values into the formula to get Then through the inverse cosine function Calculate the first included angle α.
[0080] S208. Based on the first included angle and a preset included angle threshold, determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer.
[0081] Specifically, when the first included angle is less than the preset included angle threshold, it can be determined that the pointing direction of the panoramic camera is consistent with the face orientation of the photographer. When the first included angle is not less than the preset included angle threshold, it can be determined that the pointing direction of the panoramic camera is inconsistent with the face orientation of the photographer.
[0082] In this embodiment, in practical applications, the shooting scene is complex and changeable, and the positions and postures of the shooter and the camera may change at any time. This solution uses inertial sensor data and face detection technology to be able to track the state changes of the shooter and the camera in real time, continuously calculate the included angle and make a consistency judgment, without being affected by scene changes, and always maintain a high judgment accuracy, thus ensuring the accuracy of subsequent target recognition.
[0083] In one embodiment, as Figure 3 shown, establishing a face coordinate system based on the face orientation and the direction of gravity indicated by the inertial sensor data includes:
[0084] S302, determining a first coordinate axis according to the face orientation, determining a second coordinate axis according to the direction of gravity, and determining a third coordinate axis by determining a straight line perpendicular to the first coordinate axis and the second coordinate axis.
[0085] S304, taking the face as the origin, and establishing a face coordinate system based on the first coordinate axis, the second coordinate axis, and the third coordinate axis.
[0086] Among them, the first coordinate axis is usually the coordinate axis determined according to the face orientation, which represents a main reference direction of the face in the horizontal direction. Generally speaking, the projection direction of the direction directly in front of the face on the horizontal plane is used as the direction of the first coordinate axis to describe the relevant position and angle information of the face in the horizontal direction. The direction of gravity is generally the direction of the action of the earth's gravity, which is always vertically downward in the real world. In this technical solution, it is an important reference for determining another coordinate axis of the face coordinate system, and the direction of gravity can be measured by an inertial sensor (such as an accelerometer) in the device. The second coordinate axis is usually the coordinate axis determined according to the direction of gravity. Usually, the projection direction of the direction of gravity in the direction perpendicular to the plane where the first coordinate axis is located is used as the direction of the second coordinate axis. It is mainly used to describe the information related to the vertical direction and assist in constructing a three-dimensional face coordinate system. The third coordinate axis is usually the coordinate axis determined by a straight line perpendicular to the first coordinate axis and the second coordinate axis. In three-dimensional space, according to the right-hand rule, when the first coordinate axis and the second coordinate axis are determined, the direction of the third coordinate axis is also uniquely determined. It forms a complete three-dimensional rectangular coordinate system with the first two coordinate axes to comprehensively describe the position and direction information in the face coordinate system. The face coordinate system is a three-dimensional coordinate system established with the face as the origin, and the position and direction in space are defined by the above three coordinate axes.
[0087] Specifically, the positive direction of the projection of the direction directly in front of the human face on the horizontal plane can be used as the positive direction of the first coordinate axis. The gravity acceleration data is obtained by using the accelerometer in the built-in inertial measurement unit (IMU) of the device. The accelerometer measures the acceleration of the device on three axes (assumed to be the x, y, and z axes). When the device is stationary or the motion state is relatively stable, this acceleration data contains information about the gravity acceleration. The data obtained by the accelerometer is processed to remove possible other acceleration interferences (such as the acceleration generated by the device's motion) to obtain the gravity acceleration vector. The second coordinate axis is determined based on the gravity acceleration vector. After determining the first coordinate axis and the second coordinate axis, the third coordinate axis can be determined according to the first coordinate axis, the second coordinate axis, and the right-hand rule. With the human face as the origin, a face coordinate system is established according to the first coordinate axis, the second coordinate axis, and the third coordinate axis.
[0088] In this embodiment, the coordinate axes are determined using the gravity direction, making the coordinate system have a certain stability. In a complex motion environment, even if the device's attitude changes, the gravity direction is always a relatively stable reference. By combining the face orientation and the gravity direction, the system can accurately establish a face coordinate system in different device postures, reducing errors caused by environmental interference.
[0089] In one embodiment, as Figure 4 shown, determining whether the pointing direction of the panoramic camera and the face orientation of the photographer are consistent based on the panoramic image obtained by the panoramic camera and the position of the photographer in space includes:
[0090] S402, determine the inertial sensor data currently corresponding to the panoramic image, and establish a camera coordinate system based on the gravity direction indicated by the inertial sensor data.
[0091] Among them, the camera coordinate system refers to the coordinate system established using the panoramic camera, and the origin of the camera coordinate system is the optical center of the camera. In a common right-hand coordinate system, the directions of its coordinate axes have specific definitions. Among them, the Z-axis coincides with the optical axis of the camera, and the positive direction points in front of the camera, that is, the direction of light propagation; the X-axis is horizontal to the right; the Y-axis is vertical upward. These three mutually perpendicular coordinate axes form a three-dimensional rectangular coordinate system. In this coordinate system, any point in space can be represented by a set of coordinate values (X, Y, Z) to indicate its position relative to the camera.
[0092] S404, project the panoramic image according to the camera coordinate system, determine the face orientation in the panoramic image, and calculate the Yaw attitude angle according to the face orientation.
[0093] Among them, the Yaw attitude angle can be the rotation angle used to describe the rotation of an object around the vertical axis (usually corresponding to the Z-axis of the camera coordinate system). In face pose estimation, the Yaw attitude angle represents the left-right rotation angle of the face around the axis perpendicular to the ground. A positive angle usually indicates that the face rotates to the right, and a negative angle indicates that the face rotates to the left.
[0094] Specifically, obtain the internal and external parameters of the camera. The internal parameters include focal length, principal point position, lens distortion parameters, etc., which can be obtained through camera calibration. The external parameters describe the position and attitude of the camera in the world coordinate system, usually represented by a rotation matrix and a translation vector. According to the parameters of the camera, perform a projection transformation on each pixel point in the panoramic image. For a pixel point (u, v) in the panoramic image, through the imaging model and projection formula of the camera, convert it into the three-dimensional coordinates (X, Y, Z) in the camera coordinate system. Use a face detection algorithm to detect faces in the projected panoramic image. A cascade classifier based on Haar features, a deep learning-based face detection network (such as MTCNN, YOLO, etc.) can be used. Thus, locate the position and approximate contour of the face. Extract features according to the position and approximate contour. According to the extracted face features, estimate the orientation of the face. Represent the face orientation as a direction vector in the camera coordinate system. Determine the Yaw attitude angle according to the projection of the direction vector in the XY plane of the camera coordinate system.
[0095] S406, based on the Yaw attitude angle and a preset included angle threshold, determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer.
[0096] Specifically, compare the Yaw angle of the pointing direction of the panoramic camera with the Yaw attitude angle of the face orientation of the photographer, and calculate the difference between them. When the difference is less than the preset included angle threshold, it can be determined that the pointing direction of the panoramic camera is consistent with the face orientation of the photographer. When the difference is not less than the preset included angle threshold, it can be determined that the pointing direction of the panoramic camera is inconsistent with the face orientation of the photographer.
[0097] In this embodiment, by obtaining the inertial sensor data corresponding to the panoramic image and establishing a camera coordinate system based on the indicated gravity direction, a stable and reliable reference system is provided for subsequent analysis. The gravity direction is relatively stable. The coordinate system established based on it can accurately reflect the spatial position and attitude of the camera, making the judgment of the pointing direction of the panoramic camera and the face orientation of the photographer more accurate. Calculate the Yaw attitude angle of the face orientation and compare it with the preset included angle threshold, realizing a quantitative judgment of the direction consistency between the two. This quantitative analysis method avoids the error of subjective judgment and can accurately determine the direction relationship between the camera and the face in various complex scenarios.
[0098] In one embodiment, as Figure 5As shown, establishing a camera coordinate system based on the gravity direction indicated by the inertial sensor data includes:
[0099] S502. Determine the first coordinate axis according to the gravity direction, determine the second coordinate axis according to the horizontal direction perpendicular to the gravity direction, and determine the third coordinate axis according to the gravity direction and the horizontal direction.
[0100] S504. With the panoramic camera as the origin, establish a camera coordinate system according to the first coordinate axis, the second coordinate axis, and the third coordinate axis.
[0101] Specifically, the direction of the gravity direction or the opposite direction of the gravity direction can be used as the direction of the first coordinate axis to determine the first coordinate axis. Select a suitable direction in the plane perpendicular to the gravity direction as the horizontal direction. Determine the second coordinate axis according to the horizontal direction. According to the right-hand rule, and use the first coordinate axis and the second coordinate axis to determine the third coordinate axis. With the position of the panoramic camera as the origin, and using the direction vectors of the first coordinate axis, the second coordinate axis, and the third coordinate axis as the directions of the x-axis, y-axis, and z-axis respectively, establish a camera coordinate system.
[0102] In one embodiment, as Figure 6 shown, determining a search area in the panoramic image based on the pointing direction of the panoramic camera includes:
[0103] S602. Project the panoramic image to obtain a planar panoramic image.
[0104] S604. Based on the pointing direction of the panoramic camera and the planar panoramic image, determine an initial pointing point in the planar panoramic image.
[0105] S606. Based on the initial pointing point and a preset search range, determine the search area in the planar panoramic image.
[0106] Among them, the planar panoramic image is usually a panoramic image in a two-dimensional planar form obtained after a projection operation. It shows the information of the original 360-degree scene on the plane. Although it is two-dimensional, it still retains the perspective characteristics of the panorama. The initial pointing point can be a specific point determined in the planar panoramic image according to the pointing direction of the panoramic camera. This point represents the corresponding position of the camera pointing on the planar panoramic image and is the key reference point for subsequent determination of the search area.
[0107] Specifically, select a suitable projection method according to actual needs and application scenarios, such as equirectangular projection (ERP). ERP projection is to transform the points on the sphere to the cylinder surface according to a certain mapping relationship and then unfold them into a planar image. For each point in the panoramic image, according to its longitude and latitude coordinates on the sphere (assumed to be θ and , calculate the coordinates (x, y) on the planar image through the corresponding projection formula. According to the selected projection method, convert the three-dimensional coordinates of the camera pointing direction into the two-dimensional coordinates of the planar panoramic image. For example, in the equidistant cylindrical projection, according to the longitude and latitude information of the camera pointing, use the projection formula to calculate the corresponding x and y coordinate values on the planar panoramic image, and the point corresponding to this coordinate is the initial pointing point. According to the specific application requirements, preset the shape (such as rectangle, circle, etc.) and size parameters of the search range. For example, if the preset search range is a rectangle, it is necessary to determine the length and width of the rectangle; if it is a circle, it is necessary to determine the radius. With the initial pointing point as the center, calculate the boundary coordinates of the search area in the planar panoramic image according to the preset search range shape and size. Determine the search area in the planar panoramic image through the boundary coordinates.
[0108] In this embodiment, after projecting the panoramic image into a planar panoramic image, determine the initial pointing point based on the camera pointing direction, and delimit the search area with this and the preset search range, which can significantly narrow the processing range of subsequent operations such as image analysis and target detection. The panoramic image contains 360-degree omnidirectional information, and the data volume is huge. Directly processing all information is time-consuming and laborious. By determining the search area, only process the specific area, reduce the computational amount, and improve the processing speed. Delimiting the search area avoids processing a large amount of irrelevant data and reduces data redundancy. This not only saves computing resources but also enables the algorithm to focus more on the content of interest and improves the operation efficiency of the algorithm.
[0109] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0110] Based on the same inventive concept, the embodiments of the present disclosure also provide a shooting target recognition device for implementing the shooting target recognition method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the shooting target recognition device provided below can refer to the limitations on the shooting target recognition method in the above text, and will not be repeated here.
[0111] In one embodiment, asFigure 7 As shown, a shooting target recognition device 700 is provided, including: a panoramic image acquisition module 702, a face orientation detection module 704, a search area determination module 706, and a target recognition adjustment module 708, where:
[0112] The panoramic image acquisition module 702 is used to acquire a panoramic image obtained by the shooter through a panoramic camera;
[0113] The face orientation detection module 704 is used to determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the shooter based on the panoramic image obtained by the panoramic camera and the position of the shooter in space;
[0114] The search area determination module 706 is used to, in response to the pointing direction of the panoramic camera being consistent with the face orientation of the shooter, determine a search area in the panoramic image based on the pointing direction of the panoramic camera;
[0115] The target recognition adjustment module 708 is used to determine the recognition target of the current panoramic image based on the search area and place the recognition target in the central area of the panoramic image.
[0116] In an embodiment of the device, the face orientation detection module 704 includes:
[0117] A coordinate system establishment module is used to determine the inertial sensor data corresponding to the current panoramic image, determine the face of the shooter in the panoramic image, and establish a face coordinate system based on the face orientation and the gravity direction indicated by the inertial sensor data.
[0118] A face pose determination module is used to project the panoramic camera into the face coordinate system and determine a first pose of the panoramic camera in the face coordinate system relative to the face.
[0119] An included angle determination module is used to determine the pointing direction of the panoramic camera in the first pose, and determine a first included angle between the face orientation and the pointing direction according to the planar projection of the pointing direction in the face coordinate system.
[0120] An orientation detection module is used to determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the shooter based on the first included angle and a preset included angle threshold.
[0121] In one embodiment of the device, the coordinate system establishment module is further configured to determine a first coordinate axis according to the face orientation, determine a second coordinate axis according to the gravity direction, and determine a third coordinate axis by determining a straight line perpendicular to the first coordinate axis and the second coordinate axis; with the face as the origin, a face coordinate system is established based on the first coordinate axis, the second coordinate axis, and the third coordinate axis.
[0122] In one embodiment of the device, the face orientation detection module 704 includes:
[0123] The camera coordinate system establishment module is configured to determine the inertial sensor data currently corresponding to the panoramic image, and establish a camera coordinate system based on the gravity direction indicated by the inertial sensor data.
[0124] The attitude angle calculation module is configured to project the panoramic image according to the camera coordinate system, determine the face orientation in the panoramic image, and calculate the Yaw attitude angle according to the face orientation.
[0125] The orientation detection module is further configured to determine whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer based on the Yaw attitude angle and a preset included angle threshold.
[0126] In one embodiment of the device, the camera coordinate system establishment module is further configured to determine a first coordinate axis according to the gravity direction, determine a second coordinate axis according to the horizontal direction perpendicular to the gravity direction, and determine a third coordinate axis according to the gravity direction and the horizontal direction; with the panoramic camera as the origin, a camera coordinate system is established according to the first coordinate axis, the second coordinate axis, and the third coordinate axis.
[0127] In one embodiment of the device, the search area determination module 706 is further configured to project the panoramic image to obtain a planar panoramic image; determine an initial pointing point in the planar panoramic image based on the pointing direction of the panoramic camera and the planar panoramic image; and determine a search area in the planar panoramic image based on the initial pointing point and a preset search range.
[0128] Each module in the above shooting target recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0129] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for recognizing a shooting target. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0130] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the computer device to which the solution of the present disclosure is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0131] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0132] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0133] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0134] It should be noted that the panoramic images obtained by the panoramic camera involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present disclosure can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present disclosure can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present disclosure can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0137] The above-described embodiments only represent several implementation manners of the present disclosure. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present disclosure, several modifications and improvements can still be made, and these all belong to the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the appended claims.
Claims
1. A method for shooting target recognition, characterized in that, The method includes: Obtaining a panoramic image captured by a panoramic camera by a photographer; Based on the panoramic image captured by the panoramic camera and the position of the photographer in space, determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer; In response to the pointing direction of the panoramic camera being consistent with the face orientation of the photographer, determining a search area in the panoramic image based on the pointing direction of the panoramic camera; Determining an identification target of the current panoramic image based on the search area, and placing the identification target in the central area of the panoramic image.
2. The method according to claim 1, wherein The determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer based on the panoramic image captured by the panoramic camera and the position of the photographer in space includes: Determining the inertial sensor data currently corresponding to the panoramic image, determining the face of the photographer in the panoramic image, and establishing a face coordinate system based on the face orientation and the gravity direction indicated by the inertial sensor data; Projecting the panoramic camera into the face coordinate system, and determining a first pose of the panoramic camera relative to the face in the face coordinate system; Determining the pointing direction of the panoramic camera in the first pose, and determining a first angle between the face orientation and the pointing direction according to the planar projection of the pointing direction in the face coordinate system; Based on the first angle and a preset angle threshold, determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer.
3. The method according to claim 2, wherein The establishing a face coordinate system based on the face orientation and the gravity direction indicated by the inertial sensor data includes: Determining a first coordinate axis according to the face orientation, determining a second coordinate axis according to the gravity direction, and determining a third coordinate axis by determining a straight line perpendicular to the first coordinate axis and the second coordinate axis; Taking the face as the origin, establishing a face coordinate system based on the first coordinate axis, the second coordinate axis, and the third coordinate axis.
4. The method according to claim 1, wherein The determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer based on the panoramic image captured by the panoramic camera and the position of the photographer in space includes: Determining the inertial sensor data currently corresponding to the panoramic image, and establishing a camera coordinate system based on the gravity direction indicated by the inertial sensor data; Projecting the panoramic image according to the camera coordinate system, determining the face orientation in the panoramic image, and calculating a Yaw pose angle according to the face orientation; Based on the Yaw pose angle and a preset angle threshold, determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer.
5. The method according to claim 4, wherein The establishing a camera coordinate system based on the gravity direction indicated by the inertial sensor data includes: Determining a first coordinate axis according to the gravity direction, determining a second coordinate axis according to the horizontal direction perpendicular to the gravity direction, and determining a third coordinate axis according to the gravity direction and the horizontal direction; Taking the panoramic camera as the origin, establishing a camera coordinate system according to the first coordinate axis, the second coordinate axis, and the third coordinate axis.
6. The method according to claim 1, characterized in that Determining a search area in the panoramic image based on the pointing direction of the panoramic camera includes: Projecting the panoramic image to obtain a planar panoramic image; Determining an initial pointing point in the planar panoramic image based on the pointing direction of the panoramic camera and the planar panoramic image; Determining a search area in the planar panoramic image based on the initial pointing point and a preset search range.
7. A shooting target recognition device, characterized in that, The apparatus includes: A panoramic image acquisition module for acquiring a panoramic image captured by a panoramic camera by a photographer; A face orientation detection module for determining whether the pointing direction of the panoramic camera is consistent with the face orientation of the photographer based on the panoramic image captured by the panoramic camera and the position of the photographer in space; A search area determination module for determining a search area in the panoramic image based on the pointing direction of the panoramic camera in response to the pointing direction of the panoramic camera being consistent with the face orientation of the photographer; A target recognition adjustment module for determining an identification target of the current panoramic image based on the search area and placing the identification target in the central area of the panoramic image.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.