Shooting method and wearable device
By performing multi-dimensional visual feature analysis and dynamic threshold adjustment on the preview frames of wearable devices, the problems of low shooting quality and operational complexity caused by users' inability to preview images are solved, and flexible shooting direction adjustment and improved user experience are achieved.
Patent Information
- Application Number
- CN202510888680.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
AI Technical Summary
During the shooting process, wearable devices lack a display screen, which prevents users from previewing the images in real time and making it difficult to control the shooting quality. This results in low usability of the captured materials. The fixed shooting direction cannot adapt to complex scenes, and the cumbersome switching method affects the user experience.
By performing multi-dimensional visual feature analysis on the preview frames in the target preview stream collected by wearable devices, integrating features such as the target aspect ratio, edge density ratio, and image scene category, dynamically adjusting the shooting direction, and optimizing the shooting direction switching using a dynamic threshold strategy and a time domain smoothing filtering algorithm, providing directional prompt voice guidance.
It realizes adaptive adjustment of shooting direction without preview conditions, improves shooting quality and user experience, and reduces misjudgment and operation complexity.
Smart Images

Figure CN120751243A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of photographing technology, and in particular to a photographing method and a wearable device. Background Art
[0002] With the development of technology, wearable devices are rapidly gaining popularity due to their portability and adaptability to multiple scenarios. Wearable devices with integrated cameras are widely used in sports recording, security monitoring and other fields due to their advantages such as hands-free shooting and immersive viewing angles.
[0003] During the filming process, wearable devices lack a display screen, preventing users from previewing the footage in real time. This makes it difficult to control the quality of the footage, and the usability of the captured footage is low. Related technologies use fixed shooting directions (such as horizontal or vertical), resulting in illogical composition and an inability to adapt to complex scenes. Alternatively, wearable devices can be wirelessly controlled to switch shooting directions via mobile phones or other terminal devices, which is cumbersome and disrupts the continuity of the shot, affecting the user experience.
[0004] Therefore, in order to improve the quality of wearable device footage, a flexible shooting method is urgently needed. Summary of the Invention
[0005] The present application provides a shooting method and a wearable device that can adaptively adjust the shooting direction of a shooting target.
[0006] In a first aspect, a shooting method is provided for use with a wearable device, the wearable device including a camera. The shooting method includes: determining a shooting target; obtaining a preview frame corresponding to the shooting target, the preview frame being a single frame of image data in a target preview stream captured by the camera, and the preview frame not being displayed during the capture process; fusing preset visual features of the preview frame to obtain a decision score for the preview frame, wherein the preset visual features include at least two of a target aspect ratio, an edge density ratio, and an image scene category, wherein the target aspect ratio represents the aspect ratio of the shooting target in the preview frame, the edge density ratio represents the ratio of the total vertical edge intensity to the total horizontal edge intensity in the preview frame, and the image scene category represents the scene category of the content corresponding to the preview frame; and analyzing the decision score based on a dynamic threshold strategy, wherein the dynamic threshold strategy dynamically adjusts the decision threshold according to a motion blur index.
[0007] In scenarios where the user cannot preview the captured image, a multi-dimensional visual feature analysis is performed on the preview frames corresponding to the captured target in the target preview stream collected by the wearable device. After fusing the visual features of each dimension, a decision score is obtained. The decision score is then analyzed using a dynamic threshold strategy. In scenarios where the captured image cannot be previewed, the shooting direction of the captured target is adjusted, thereby achieving adaptive adjustment of the shooting direction of the captured target.
[0008] In one possible implementation, the decision threshold includes a first threshold and a second threshold, and the first threshold is greater than the second threshold. The decision score is analyzed based on a dynamic threshold strategy, and the shooting direction of the shooting target is adjusted, including: when the decision scores corresponding to the preview frames of M consecutive frames are all greater than the first threshold, the shooting direction of the shooting target is adjusted to a vertical mode; when the decision scores corresponding to the preview frames of M consecutive frames are all less than the second threshold, the shooting direction of the shooting target is adjusted to a horizontal mode; wherein M is a positive integer greater than 1.
[0009] Furthermore, if the decision scores corresponding to the preview frames of M consecutive frames meet the first condition or the second condition, the shooting direction is maintained; the first condition includes: the decision scores corresponding to the preview frames of M consecutive frames are all within the numerical range between the first threshold and the second threshold; the second condition includes: the decision scores corresponding to the preview frames of M consecutive frames are not all greater than the first threshold, not all less than the second threshold, and not all within the numerical range between the first threshold and the second threshold.
[0010] In an embodiment of the present application, under the condition that the decision threshold includes a first threshold and a second threshold, and the first threshold is greater than the second threshold, a time domain smoothing filtering algorithm is introduced to process multiple preview frames of a time series, thereby reducing misjudgment of instantaneous jitter and improving the accuracy of shooting direction switching.
[0011] In a possible implementation, preset visual features including a target aspect ratio, an edge density ratio, and an image scene category are fused to obtain a decision score for the preview frame, including: analyzing the preview frame based on a segmentation algorithm to obtain an outline of the captured target and determining the target aspect ratio of the preview frame; analyzing the preview frame based on an edge detection algorithm and directional density to obtain an edge density ratio of the preview frame; determining the image scene category corresponding to the preview frame based on a preset image classification model, and determining a semantic scene weight value of the preview frame based on the image scene category, a preset mapping relationship between the preset scene category and the horizontal weight addition and the vertical weight addition, and a classification confidence; performing feature fusion weighted summation on the target aspect ratio, the edge density ratio, and the semantic scene weight value of the preview frame to obtain a decision score.
[0012] The preview frame is analyzed based on a segmentation algorithm to obtain the outline of the shooting target, and the target aspect ratio of the preview frame is determined, including: separating the outline of the shooting target from the preview frame based on the segmentation algorithm; dividing the height of the outline by the width of the outline to obtain a first aspect ratio; when there is a single shooting target in the preview frame, determining the target aspect ratio of the preview frame based on the first aspect ratio of the shooting target; when there are multiple shooting targets in the preview frame, determining the target aspect ratio of the preview frame based on the first aspect ratio of each shooting target and a target saliency weight corresponding to each shooting target, where the target saliency weight is used to represent a quantified weight of the visual saliency of the shooting target in the preview frame, where the visual saliency includes at least size and position.
[0013] In addition, the preview frame is analyzed based on the edge detection algorithm and directional density to obtain the edge density ratio of the preview frame, including: determining the horizontal gradient and vertical gradient of the preview frame based on the Sobel operator; dividing the preview frame into multiple preset grids, and determining the total horizontal edge intensity and the total vertical edge intensity of the preset grids; determining the density ratio based on the total horizontal edge intensity and the total vertical edge intensity of each preset grid; performing segmented conversion based on the density ratio to obtain the edge density ratio of the preview frame.
[0014] Under the condition that the computing power of the device is relatively sufficient, the embodiments of the present application increase the number of image features covered in the preset visual features, that is, the preset visual features include the target aspect ratio, edge density ratio and image scene category, improve the completeness of the features presented in the preview frame, and improve the generalization ability and robustness of the shooting method.
[0015] In one possible implementation, the wearable device also includes a target speaker. After determining the shooting target, it also includes: analyzing the preview frame based on the target detection algorithm to obtain the truncation degree of the shooting target; if the truncation degree of the shooting target is greater than a preset safety distance threshold, determining that the preview frame is in a truncated state; when the preview frame is in the truncated state, generating a direction prompt voice, and outputting the direction prompt voice through the target speaker.
[0016] In addition, the target speaker includes a first speaker and a second speaker, and the first speaker and the second speaker are respectively arranged at different positions of the wearable device, and the direction prompt voice is played through the target speaker, including: based on the direction information carried by the direction prompt voice; when the direction information is a first direction, determine the first speaker corresponding to the first direction, and play the direction prompt voice through the first speaker; when the direction information is a second direction, determine the second speaker corresponding to the second direction, and play the direction prompt voice through the second speaker.
[0017] In the scene of shooting a target, if the target is incomplete, during the horizontal and vertical switching process, the user can be guided to adjust the shooting angle through the cropping analysis of the target to improve the completeness of the target in the shooting result.
[0018] In a second aspect, a photographing device is provided, comprising a unit for executing any one of the methods in the first aspect. The device may be a terminal device or a chip within the terminal device.
[0019] In a third aspect, a wearable device is provided, comprising a processor and a camera, wherein the processor is configured to execute any one of the shooting methods of the first aspect.
[0020] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a wearable device, the wearable device executes any one of the shooting methods in the first aspect.
[0021] In a fifth aspect, a computer program product is provided, which includes a computer program. When the computer program is executed by a wearable device, the wearable device executes any one of the shooting methods of the first aspect.
[0022] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 shows a schematic diagram of the mechanical structure of smart glasses;
[0024] Figure 2 A schematic diagram showing a picture taken by a user wearing smart glasses while hiking in a mountain;
[0025] Figure 3 A schematic diagram showing a desired shot of a user wearing smart glasses while hiking in a mountain;
[0026] Figure 4 A hardware structure diagram of a wearable device provided in an embodiment of the present application is shown;
[0027] Figure 5 A schematic diagram showing a process of a photographing method provided in an embodiment of the present application is shown;
[0028] Figure 6 A schematic diagram showing a flow chart of another photographing method provided in an embodiment of the present application is shown;
[0029] Figure 7 A schematic diagram showing a flow chart of another photographing method provided in an embodiment of the present application is shown;
[0030] Figure 8 A flow chart corresponding to a shooting method provided in an embodiment of the present application is shown;
[0031] Figure 9 A schematic diagram showing a flow chart of another photographing method provided in an embodiment of the present application is shown;
[0032] Figure 10 A schematic diagram of a shooting device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0034] It should be understood that the “multiple” mentioned in this application refers to two or more. In the description of this application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solution of this application, words such as “first” and “second” are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as “first” and “second” do not limit the quantity and execution order, and words such as “first” and “second” do not necessarily limit them to be different.
[0035] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in different places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. Furthermore, the terms "including," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0036] Before explaining in detail the shooting method and wearable device provided in the embodiments of the present application, the application scenarios and related technologies of the shooting method and wearable device are first explained.
[0037] Wearable devices refer to portable electronic devices that are worn directly on the human body or integrated into clothing or accessories. These devices can realize functions such as data collection, information processing and human-computer interaction through direct contact with the human body or integration into clothing.
[0038] With the development of technology, wearable devices are rapidly gaining popularity due to their portability and adaptability to multiple scenarios. Wearable devices with integrated cameras offer significant advantages, supporting hands-free shooting and providing an immersive perspective. Consequently, they are widely used in sports recording, security monitoring, and other fields.
[0039] Take wearable devices such as smart glasses as an example. Figure 1 The mechanical structure diagram of the smart glasses is shown in FIG. Figure 1 As shown in FIG, the main structure of the smart glasses consists of temples 11, frames 12 and lenses 13, and is equipped with functional components such as cameras, physical buttons, microphones, and speakers to provide intelligent services for users. Among them, the number of cameras, physical buttons, microphones, and speakers can be one or more, such as Figure 1 As shown, there are cameras 21 and 22, physical buttons 31 and 32, and microphones 41 and 42.
[0040] like Figure 1 As shown, as an example, the camera 21 and the camera 22 are integrated into the frame 12; the physical button 31 and the physical button 32 are distributed on the temple 11; the speaker 41 and the speaker 42 are respectively arranged on the first temple and the second temple 11; the microphone is integrated into the frame 12; the various functional components in the smart glasses can be flexibly arranged according to design requirements, and there is no restriction on the integration position.
[0041] In outdoor adventure scenarios, when users use smart glasses to record hiking and mountaineering, they cannot preview the captured images in real time due to the lack of a display screen, making it difficult to judge whether the composition of the captured images is reasonable. The captured materials often result in problems such as distorted picture proportions and missing targets due to improper shooting directions.
[0042] When it comes to shooting direction control, wearable devices such as smart glasses, if they use a fixed shooting direction (such as horizontal or vertical shooting), will find it difficult to flexibly adapt to the complex and changing scenes during mountain climbing, resulting in poor composition of the captured images (such as excessive white space in the image, mountain distortion, etc.); if the shooting direction of the wearable device is switched wirelessly through terminal devices such as mobile phones, it will be inconvenient to operate during climbing and will also interrupt the continuity of the shooting, which may result in missing precious images and affecting the user experience.
[0043] In addition, currently, mobile phones, digital cameras and other devices used for shooting are widely used, and users are accustomed to shooting methods with previews. For the shooting methods without previews corresponding to wearable devices, users need to go through learning guidance and practice to improve their experience of shooting with wearable devices.
[0044] Figure 2 A schematic diagram showing a picture taken by a user wearing smart glasses while hiking. Figure 3 The figure shows a user wearing smart glasses and expecting to take pictures while hiking. The user wants to use smart glasses to record a vertical picture of a steep mountain wall and a magnificent sea of clouds. Figure 3 As shown; However, because the default shooting direction of the smart glasses is horizontal, the final picture is horizontal, as shown Figure 2 As shown, a large area of dark mountains occupies the screen, the sky and the mountains are out of proportion, and the visual effect is poor; if you try to take out your mobile phone and adjust the shooting direction through the application (Application, APP) in the mobile phone, it is inconvenient to operate in cold environments and may also bring safety hazards due to distraction.
[0045] For example, in an industrial inspection scenario, users (engineers) wear smart glasses to capture equipment details. Due to the inability to preview the images, problems such as improper shooting direction often occur, resulting in the captured images being unable to be retained as valid data.
[0046] In view of this, embodiments of the present application provide a shooting method and a wearable device, which are applied to a wearable device including a camera. The method comprises determining a shooting target; obtaining a preview frame corresponding to the shooting target, wherein the preview frame is a single frame of image data in a preview stream of the target captured by the camera, and the preview frame is not displayed during the capture process; fusing preset visual features of the preview frame to obtain a decision score for the preview frame; the preset visual features include at least two of the target aspect ratio, edge density ratio, and image scene category; and further analyzing the decision score based on a dynamic threshold strategy to adjust the shooting direction of the shooting target. The dynamic threshold strategy dynamically adjusts the decision threshold based on a motion blur index. In scenarios where the user cannot preview the captured image, the preview frames in the real-time preview stream captured by the wearable device are subjected to multi-dimensional visual feature analysis. After fusing the visual features of each dimension, a decision score is obtained. The decision score is further analyzed using the dynamic threshold strategy to adjust the shooting direction of the shooting target in scenarios where the user cannot preview the captured image, thereby achieving adaptive adjustment of the shooting direction of the shooting target.
[0047] The following describes the shooting method provided in the embodiment of the present application in conjunction with the accompanying drawings. The method is applied to wearable devices. The wearable devices can be smart glasses (such as AR glasses, etc.), shooting glasses, sports cameras, thumb cameras, audio and video recorders and other electronic devices. The embodiment of the present application does not impose any special restrictions on the specific technology and specific form adopted by the wearable device.
[0048] Figure 4 FIG. 1 shows a hardware structure diagram of a wearable device provided in an embodiment of the present application. For example, taking smart glasses as an example, Figure 4 As shown, the wearable device 100 may include a processor 110, a memory 120, a camera 130, a sensor module 140, a charging management module 150, a battery 151, a power management module 152, a wireless communication module 160, an audio module 170, a speaker 171, and a microphone 172; wherein the sensor module 140 may include a gyroscope sensor 141, an acceleration sensor 142, and the like.
[0049] It should be noted that Figure 4 The structure shown does not constitute a specific limitation on the wearable device 100. In other embodiments of the present application, the wearable device 100 may include Figure 4 More or fewer components than those shown, or wearable device 100 may include Figure 4 A combination of some of the components shown, or the wearable device 100 may include Figure 4 Subassemblies of some of the components shown. Figure 4 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0050] The processor 110 may include one or more processing units, such as a microcontroller unit (MCU), a central processing unit (CPU), memory, peripheral interfaces, etc., and a neural network processor (NPU). The NPU is a processor specially designed for neural network computing, which has higher efficiency and performance when processing computing tasks related to deep learning models. Different processing units can be independent devices or integrated into one or more processors.
[0051] The processor 110 may also be provided with a memory for storing instructions and data.
[0052] The memory 120 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the wearable device 100 by running the instructions stored in the memory 120.
[0053] It should be understood that the memory 120 may include a high-speed random access memory (RAM), etc. Among them, a high-speed random access memory is a computer memory. In a wearable device, RAM is a place for temporarily storing running programs and data during operation.
[0054] The gyroscope sensor 141 and the acceleration sensor 142 may constitute an inertial measurement unit (IMU) to determine the motion state, for example, the motion blur index, by measuring the angular velocity and acceleration.
[0055] The charging management module 150 is used to receive charging input from a charger, which can be a wireless charger or a wired charger.
[0056] The wireless communication module 160 is used for data transmission and interaction with terminal devices such as mobile phones, and realizes multi-dimensional function expansion through technologies such as Bluetooth and Wi-Fi.
[0057] An operating system runs on the wearable device 100, such as an Android operating system or a customized version thereof, a real-time operating system (RTOS) or an optimized version thereof. The embodiment of the present application takes RTOS as an example. It should be noted that although the embodiment of the present application is described using a wearable device operating system as an example, its basic principles are also applicable to wearable devices 100 with other operating systems.
[0058] For example, smart glasses running an RTOS system utilize technologies like pruning and quantization to create lightweight models that fit within the limited memory and computing resources of smart glasses, allowing them to run efficiently on low-power MCUs or NPUs. During runtime, the model rapidly processes preview frames captured by the camera, integrating multi-dimensional features to achieve target recognition and scene understanding. The RTOS system's efficient task scheduling ensures that inference and other real-time tasks (such as display refresh and sensor data processing) do not interfere with each other, ensuring a balanced balance between performance and real-time performance.
[0059] To facilitate a further understanding of the technical solutions in some embodiments of the present application, the following describes in detail the technical solutions of the shooting method and how the technical solutions solve the above-mentioned technical problems in conjunction with some specific embodiments and drawings. The various embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. Obviously, the described embodiments are only some of the embodiments of the present application, not all of them.
[0060] The following combination Figures 5 to 9 The shooting method provided in the embodiment of the present application is described in detail. Figure 5 A schematic diagram of a process flow of a shooting method provided by an embodiment of the present application is shown, which is applied to the above Figure 1 and Figure 4 The wearable device shown in FIG. 1 is a smart glass, which includes a camera and does not have a display screen. Figure 5 As shown, the shooting method includes the following steps:
[0061] S210: Determine the shooting target.
[0062] The shooting target can be determined based on the target detection algorithm, which analyzes the captured image in real time and determines the preset target when a high-confidence target (such as a car, a person, etc.) is identified; the shooting target can also be determined based on the recognition of the user's visual attention target; the shooting target can also be determined through voice commands or gesture commands.
[0063] It should be understood that the shooting target can be one target or multiple targets; if there are multiple targets, their types can present diverse characteristics (such as including different categories such as people, scenery, dynamic objects, etc.).
[0064] It should be understood that when the wearable device includes a microphone, the shooting target can be determined by voice instructions, specifically including the following steps (1)-(3):
[0065] (1) When the microphone captures the voice command, determine the user intention corresponding to the voice command.
[0066] After the built-in microphone of the smart glasses captures the voice command, the built-in speech recognition model of the smart glasses will convert the voice signal into text information. Then, the built-in natural language processing (NLP) model of the smart glasses will analyze the recognized text information and understand the user's intention.
[0067] It should be understood that in actual application scenarios, the user intentions corresponding to the voice commands captured by the microphone are rich and diverse, and there are many categories, including shooting operations, listening to music operations, recording operations, etc. Therefore, it is necessary to analyze the user intentions.
[0068] (2) When the user's intention is a shooting operation, the target object is parsed from the voice command to obtain the shooting target.
[0069] When the user says "Help me take a picture of XXX", the category of the user intention corresponding to the voice instruction is a shooting operation. Therefore, the target object XXX is parsed from the voice instruction, and XXX is the shooting target.
[0070] When the user says "help me take a picture of XXX and YYY", the category of user intent corresponding to the voice command is a shooting operation. Therefore, the target objects XXX and YYY are parsed from the voice command, and both XXX and YYY are shooting targets.
[0071] (3) The initial preview stream captured by the camera is identified based on the target detection algorithm. After the shooting target is identified in the initial preview stream, the preview frame corresponding to the shooting target is determined.
[0072] The camera of the smart glasses will capture a real-time video stream of the scene in front of the lens to form a preview screen (i.e., preview stream); use target detection algorithms such as Single Shot MultiBox Detector (SSD), YOLO (You Only Look Once), and Faster Region-based Convolutional Neural Network (Faster R-CNN) to identify the target to be photographed (i.e., the shooting target); after the shooting target is identified, the picture containing the shooting target is the preview frame.
[0073] Furthermore, the preview frame can be analyzed based on the target detection algorithm, and a bounding box of each captured target, i.e., a target bounding box (Bounding Box), can be generated. It should be understood that the embodiment of the present application does not limit the type of target detection algorithm.
[0074] S220: Acquire a preview frame corresponding to the shooting target. The preview frame is not displayed during the acquisition process.
[0075] A preview frame is a single frame of the image that contains the target being captured. Specifically, it refers to the single frame of image data extracted from the target preview stream when the camera is capturing the scene in real time.
[0076] When the smart glasses are shooting, the camera will capture the real-time image of the shooting target. At this time, the preview frame is a single frame in the real-time captured target preview stream.
[0077] It should be noted that since wearable devices such as smart glasses do not have display screens, they cannot present the preview stream or preview frame during shooting, nor can they display the photos or videos that have been taken, resulting in users being unable to know the shooting status through preview.
[0078] In some application scenarios, the wearable device may have a display screen, but in a real-time acquisition scenario, the display screen does not display the preview stream corresponding to the shooting, and further, does not display the preview frame corresponding to the shooting target.
[0079] S230: Fusing the preset visual features of the preview frame to obtain a decision score of the preview frame.
[0080] Preset visual features refer to a multi-dimensional set of image features. By analyzing such features, the wearable device can enhance its understanding of the preview frame and better infer the optimal shooting direction of the target based on the preview frame.
[0081] The preset visual features include at least two of the target aspect ratio, edge density ratio, and image scene category. For example, the preset visual features may include the target aspect ratio and edge density ratio; the preset visual features may include the target aspect ratio and image scene category; the preset visual features may include the edge density ratio and image scene category; and the preset visual features may include the target aspect ratio, edge density ratio, and image scene category. In other embodiments, the preset visual features may also include other image features in addition to the above features.
[0082] Among them, the target aspect ratio represents the ratio of the length to the width of the photographed target in the preview frame, the edge density ratio represents the ratio of the total vertical edge intensity to the total horizontal edge intensity in the preview frame, and the image scene category represents the scene category of the corresponding content of the preview frame (such as people, scenery, etc.).
[0083] The target aspect ratio is a fundamental element of preview frame composition, the edge density ratio is used to refine compositional details, and the image scene category optimizes composition strategies based on scene semantic information. From an analytical perspective, pre-set visual features that include this fundamental element (target aspect ratio) perform significantly better in composition analysis than feature combinations that don't.
[0084] According to the combination method of preset visual features, preset visual features can be divided into the following types: preset visual features include target aspect ratio and edge density ratio (basic and refined), taking into account the basic composition framework and detail optimization; preset visual features include target aspect ratio and image scene category (basic and scene semantics), dynamically adjusting the composition strategy through scene attributes; preset visual features include target aspect ratio, edge density ratio and image scene category at the same time, realizing full-dimensional analysis from basic composition to semantic optimization.
[0085] Fusing the preset visual features of the preview frame is a process of combining multi-dimensional image features to form a more representative synthesis and features. It is to improve the ability to understand the preview frame in the shooting scene through information complementation and other means.
[0086] Multi-dimensional feature fusion methods include weighted fusion, splicing fusion, and attention mechanism fusion. Taking the preset visual features including object aspect ratio, edge density ratio, and image scene category as an example, the object aspect ratio, edge density ratio, and image scene category are weightedly fused to obtain a decision score.
[0087] S240: Analyze the decision score based on the dynamic threshold strategy and adjust the shooting direction of the shooting target.
[0088] The dynamic threshold strategy dynamically adjusts the decision threshold based on the Motion Blur Index (MBI). By quantifying the degree of motion blur in the image, the decision threshold is adaptively adjusted to optimize the processing performance of tasks such as decision score object detection and image segmentation. A mathematical mapping relationship between the basic threshold and the Motion Blur Index can be established to dynamically calculate and determine the decision threshold, achieving adaptive adjustment of the decision threshold.
[0089] The present invention provides a shooting method and a wearable device, which are applied to a wearable device including a camera. The method includes determining a shooting target; obtaining a preview frame corresponding to the shooting target, wherein the preview frame is a single frame of image data in a preview stream of the target captured by the camera, and the preview frame is not displayed during the capture process; fusing preset visual features of the preview frame to obtain a decision score for the preview frame; the preset visual features include at least two of the target aspect ratio, edge density ratio, and image scene category; and then analyzing the decision score based on a dynamic threshold strategy to adjust the shooting direction of the shooting target. The dynamic threshold strategy dynamically adjusts the decision threshold based on a motion blur index. In scenarios where the user cannot preview the captured image, the preview frames in the real-time preview stream captured by the wearable device are subjected to multi-dimensional visual feature analysis. After fusing the visual features of each dimension, a decision score is obtained. The decision score is then analyzed using the dynamic threshold strategy to adjust the shooting direction of the shooting target in scenarios where previewing the captured image is not possible.
[0090] In order to improve the stability of the shooting direction adjustment, dual thresholds (that is, the decision threshold includes a first threshold and a second threshold, and the first threshold is greater than the second threshold) and adaptive adjustment of the thresholds can be adopted. The decision score is analyzed based on the dynamic threshold strategy to adjust the shooting direction of the shooting target.
[0091] The decision threshold can be calculated by the following formula:
[0092] T=T0±ω(1-MBI)
[0093] Where T is the decision threshold, T0 is the basic threshold, ω is the offset amplitude, and MBI is the motion blur index.
[0094] The MBI is determined by the wearable device's IMU and ranges from 0 to 1. Larger values indicate more severe image jitter or blur, causing the decision threshold to shift toward the baseline threshold, reducing the risk of false triggers. The offset amplitude controls the adjustment range of the decision threshold. When the image is stable (i.e., MBI = 0), the initial dual thresholds are restored. When the image is severely blurred (i.e., MBI = 1), the dual thresholds converge to the baseline threshold, forcing the device to maintain automatic camera orientation.
[0095] The embodiment of the present application adopts dual thresholds and adaptive adjustment of the dual thresholds to balance response speed and anti-interference ability and adapt to stability requirements in different scenarios.
[0096] Since the first threshold is greater than the second threshold, when T0=0.5 and ω=0.15, the first threshold T1=0.5+0.15(1-MBI), and the second threshold T2=0.5-0.15(1-MBI). When MBI=0, the initial first threshold T1=0.65, and the initial second threshold T2=0.35.
[0097] If the decision score corresponding to the preview frame is greater than the first threshold, it indicates that the shooting direction of the shooting target presented in the preview frame needs to be adjusted to the portrait mode for better shooting effect; if the decision score corresponding to the preview frame is less than the second threshold, it indicates that the shooting direction of the shooting target presented in the preview frame needs to be adjusted to the landscape mode for better shooting effect; if the decision score corresponding to the preview frame is less than or equal to the first threshold and greater than or equal to the second threshold, it indicates that the shooting effect of the shooting target presented in the preview frame in the landscape mode and the portrait mode is not much different.
[0098] Therefore, when the decision score corresponding to the preview frame is greater than the first threshold, the shooting direction of the shooting target can be adjusted to the portrait mode; when the decision score corresponding to the preview frame is less than the second threshold, the shooting direction of the shooting target can be adjusted to the landscape mode; when the decision score corresponding to the preview frame is less than or equal to the first threshold and greater than or equal to the second threshold, the shooting direction is maintained unchanged.
[0099] It should be understood that when analyzing the decision score based on the dynamic threshold strategy and adjusting the shooting direction of the shooting target, it is not necessary to determine the current shooting direction of the wearable device. In other words, if the current shooting direction is landscape, and the decision score is less than the second threshold, the corresponding instruction to adjust the shooting direction of the shooting target to landscape mode is also executed. In other embodiments, the current shooting direction may also be obtained first, and if the current shooting direction is inconsistent with the shooting direction determined based on the decision score, the corresponding instruction to adjust the shooting direction is executed.
[0100] In order to reduce the misjudgment of instantaneous jitter and improve the accuracy of shooting direction switching, under the condition that the above decision threshold includes the first threshold and the second threshold, and the first threshold is greater than the second threshold, a time domain smoothing filter algorithm can be introduced to process multiple preview frames of the time series. Figure 6 A schematic diagram of a process flow of another shooting method provided in an embodiment of the present application is shown. Figure 6 As shown, for Figure 5 Step 240 in the embodiment analyzes the decision score based on the dynamic threshold strategy and adjusts the shooting direction of the shooting target, including the following steps:
[0101] S241. Obtain decision scores corresponding to M consecutive preview frames in a time series, where M is a positive integer greater than 1, such as 3, 5, 10, 15, etc.
[0102] S242: When the decision scores corresponding to the preview frames of the consecutive M frames are all greater than the first threshold, the shooting direction of the shooting target is adjusted to the portrait mode.
[0103] S243: When the decision scores corresponding to the preview frames of the consecutive M frames are all smaller than the second threshold, the shooting direction of the shooting target is adjusted to a horizontal mode.
[0104] Only when the decision scores corresponding to multiple consecutive preview frames are all less than the first threshold, the shooting direction of the shooting target can be adjusted to the horizontal mode to reduce misjudgment caused by instantaneous shaking.
[0105] S244 : If, among the decision scores corresponding to the preview frames of the consecutive M frames, all the decision scores are within a numerical range between the first threshold and the second threshold, the shooting direction is maintained.
[0106] S245. If, among the decision scores corresponding to the preview frames of the consecutive M frames, not all are greater than the first threshold, not all are less than the second threshold, and not all are within the numerical range between the first threshold and the second threshold, maintain the shooting direction.
[0107] In some embodiments, the decision scores corresponding to the preview frames of the consecutive M frames may be within a numerical range between a first threshold and a second threshold as a first condition, and the decision scores corresponding to the preview frames of the consecutive M frames may be not all greater than the first threshold, not all less than the second threshold, and not all within the numerical range between the first threshold and the second threshold as a second condition. Steps 244 and 245 may be expressed as follows: if the decision scores corresponding to the preview frames of the consecutive M frames meet the first condition or the second condition, then the shooting direction is maintained;
[0108] For example, when M is 3, the first threshold T1 = 0.65, and the second threshold T2 = 0.35, the shooting direction adjustment may be determined according to the following Table 1.
[0109] Table 1 Example of decision scores for 3 consecutive preview frames
[0110]
[0111] That is, the shooting direction is adjusted only when the decision scores corresponding to multiple consecutive preview frames are all greater than the first threshold or less than the second threshold, thereby reducing misjudgment caused by instantaneous jitter and improving the accuracy of shooting direction switching.
[0112] In order to improve the completeness of the features presented in the preview frame and improve the generalization and robustness of the shooting method, the number of image features included in the preset visual features can be increased under the condition that the device computing power is sufficient. The preset visual features include the object aspect ratio, edge density ratio and image scene category. Figure 7 A schematic diagram of a process flow of another shooting method provided in an embodiment of the present application is shown. Figure 7 As shown, for Figure 5 Step 230 in the above method fuses the preset visual features of the preview frame to obtain a decision score of the preview frame, including the following steps:
[0113] S231 , analyzing the preview frame based on a segmentation algorithm to obtain the outline of the shooting target and determine the target aspect ratio of the preview frame.
[0114] Determining the target aspect ratio of the preview frame includes the following steps (1)-(4):
[0115] (1) Separate the outline of the shooting target from the preview frame based on the segmentation algorithm. It should be understood that there may be one or more shooting targets in the preview frame. For each shooting target, a segmentation algorithm such as the interactive foreground extraction algorithm based on iterated graph cuts (GrabCut) or the mask region-based convolutional neural network model (Mask R-CNN) can be used to separate the outline of the shooting target from the preview frame.
[0116] (2) Divide the height of the contour by the width of the contour to obtain the first aspect ratio.
[0117] The first aspect ratio can be expressed as: AR=H / W, where AR is the first aspect ratio, H is the height of the outline, and W is the width of the outline.
[0118] Since the number of shooting targets in the preview frame is different, the aspect ratio of the targets in the preview frame is determined differently in the case of a single shooting target and multiple shooting targets.
[0119] (3) When there is a single photographic object in the preview frame, the target aspect ratio of the preview frame is determined based on the first aspect ratio of the photographic object.
[0120] The target aspect ratio is calculated using the following formula: Waspect = AR.
[0121] (4) When there are multiple shooting targets in the preview frame, the target aspect ratio of the preview frame is determined based on the first aspect ratio of each shooting target and the target saliency weight corresponding to each shooting target.
[0122] The target saliency weight is used to represent the weight of the quantified visual saliency of the captured target in the preview frame. The visual saliency features include at least size and position factors.
[0123] The final target aspect ratio can be determined based on the first aspect ratio of each target and its corresponding target saliency weight. Alternatively, the visual saliency analysis of each target can be used to identify the central target and other non-central targets, thereby reducing the interference of other non-central targets on the target aspect ratio. In this case, the target aspect ratio can be calculated using the following formula:
[0124] Waspect=θ1AR main +θ2AR others
[0125] Where, AR main is the first aspect ratio of the central shooting target, θ1 is the target saliency weight of the central shooting target, AR others is the average of the first aspect ratios of other non-center shooting targets, and θ2 is the target saliency weight of other non-center shooting targets.
[0126] In some examples, after visual saliency analysis is performed on each target to identify the central target and other non-central targets, the target aspect ratio can be determined using the following formula:
[0127] Waspect=θ1AR main +θ2×max(AR others )
[0128] Where, max(AR others ) is the largest first aspect ratio among other non-center shooting targets.
[0129] The target aspect ratio reflects the aspect ratio characteristics of the captured object in the preview frame and is the basis for composition decisions. Therefore, when the preset visual features include the target aspect ratio, the accuracy of the preview frame composition decision is improved.
[0130] S232: Analyze the preview frame based on the edge detection algorithm and the direction density to determine the edge density ratio of the preview frame.
[0131] The edge density ratio is used to quantify the distribution of preview frame edge information and its impact on composition. Image edges contain a wealth of structural information and are key to understanding image content. By analyzing the horizontal and vertical edge distribution, we can determine the image's visual balance and choose a more appropriate horizontal or vertical composition. Furthermore, this weight can be dynamically adjusted based on image content, making composition decisions more flexible and adaptable.
[0132] Determining the edge density ratio of the preview frame includes the following steps (1)-(4):
[0133] (1) Determine the horizontal gradient and vertical gradient of the preview frame based on the Sobel operator.
[0134] The edge density map (two-dimensional matrix) is extracted by the Sobel operator, and each pixel value represents the probability density of the edge intensity at that location in a specific direction, providing directional distribution characteristics and density statistical properties.
[0135] The calculation of the edge density map comprehensively considers global and local information and focuses on spatial continuity: on the one hand, statistics are performed on each pixel position of the entire image; on the other hand, the focus is on the edge-significant area, and a higher density value is assigned to the part where the gradient exceeds the preset threshold; at the same time, the kernel function is used to diffuse the discrete edge points into a continuous density field, thereby achieving comprehensive quantification of the image edge distribution.
[0136] Among them, the horizontal gradient corresponding to the preview frame can be expressed as:
[0137] Gx=cv2.Sobel(img,cv2.CV_16SC1,1,0,ksize=3)
[0138] In the software expression, cv2.Sobel is a function in the OpenCV library used to calculate the Sobel operator, which is used to calculate the gradient of an image. In this function, img represents the input image data (i.e., the preview frame) and is the object of the Sobel operation. cv2.CV_16SC1 is the data type flag, which indicates that the output image is a 16-bit signed integer and is used to store the gradient calculation results to avoid problems such as numerical overflow. 1,0 indicates that the gradient in the horizontal direction (X direction) is calculated, and ksize=3 indicates that the kernel size of the Sobel operator is 3×3.
[0139] And the vertical gradient is expressed as:
[0140] Gy=cv2.Sobel(img,cv2.CV_16SC1,0,1,ksize=3)
[0141] In the software expression, 0,1 means calculating the gradient in the vertical direction (Y direction).
[0142] The Sobel operator is an edge detection algorithm. The horizontal gradient and the vertical gradient can also be determined by other edge detection algorithms such as the Prewitt operator or other convolution kernels.
[0143] (2) Divide the preview frame into a plurality of preset grids, and determine the total horizontal edge strength and the total vertical edge strength of the preset grids.
[0144] The preview frame may be divided based on a preset size, or the preset frame may be divided into a preset number of grids (ie, preset grids).
[0145] Exemplarily, the preview frame is divided into 8 equal parts horizontally and vertically, forming 8×8 preset grids, a total of 64 grids. For each preset grid.
[0146] The horizontal gradient Gx is summed in the preset grid cells (i.e., sum) to obtain the total horizontal edge intensity of the preset grid, where the total horizontal edge intensity is expressed as:
[0147] horizontal_edges[i,j]=sum(Gx[cell])
[0148] The vertical gradient Gy is summed in the preset grid cells to obtain the total vertical edge intensity of the preset grid, where the total vertical edge intensity is expressed as:
[0149] vertical_edges[i,j]=sum(Gy[cell])
[0150] (3) Determine the density ratio based on the total horizontal edge intensity and the total vertical edge intensity of each preset grid.
[0151] The density ratio is expressed by the following formula:
[0152]
[0153] Where EDR is the density ratio.
[0154] (4) Perform segmented conversion based on the density ratio to obtain the edge density ratio of the preview frame.
[0155]
[0156] Where Wedge is the edge density ratio, EDR ≥ 1 represents the vertical direction, and EDR < 1 represents the horizontal direction.
[0157] S233: Determine the image scene category corresponding to the preview frame based on a preset image classification model.
[0158] You can use a preset image classification model, such as the lightweight EfficientNet-Lite0 model, to identify the scene type in the preview frame. To improve the efficiency of scene analysis for preview frames, you can retain only the top N scene categories, where N is 5, 6, etc. For example, the top 5 scene categories include people, scenery, documents, food, and architecture. The document category represents scenes of documents, books, and other objects.
[0159] S234 , analyzing the image scene category, the preset mapping relationship between the scene category and the horizontal weight addition and the vertical weight addition, and the classification confidence to determine the semantic scene weight value of the preview frame.
[0160] It's important to note that the semantic scene weight reflects the impact of semantic information in the preview frame on the composition. In terms of scene understanding, it provides semantic support for composition by identifying the scene and object categories in the preview frame. For example, for landscape images depicting natural scenery or city panoramas, horizontal composition better presents the image information because their content often extends horizontally. For portrait images featuring people, vertical composition can better highlight the person's figure and facial details, focusing the visual center.
[0161] Semantic scene weights can be combined with common compositional rules and scene characteristics to assist in making informed compositional decisions. In photography, different scene categories often correspond to specific compositional orientation preferences. For example, when photographing large scenes like mountains and coastlines, a horizontal composition aligns with the human eye's horizontal field of vision and conveys a sense of grandeur. When photographing close-ups of towering buildings or flowers, a vertical composition better reflects the vertical extension of the object.
[0162] The preset mapping relationship between the preset scene category and the horizontal weight addition and the vertical weight addition is a preset mapping relationship. The following Table 2 is an example of the preset relationship.
[0163] Table 2 Preset mapping relationship between preset scene categories and horizontal weight additions and vertical weight additions
[0164] Scene Category Vertical weight bonus Horizontal weight bonus figure +0.25 -0.15 landscape -0.10 +0.20 document +0.30 -0.25 food +0.15 -0.10 architecture -0.05 +0.18
[0165] Based on the classification confidence, the nonlinear mapping is analyzed to determine the semantic scene weight value of the preview frame as follows:
[0166] Wscene=baseweight+Δ×confidence
[0167] Where Wscene is the semantic scene weight value, baseweight is the base weight, Δ is the weight addition corresponding to the scene category (vertical weight addition or horizontal weight addition), and confidence is the classification confidence output by the preset image classification model.
[0168] The default value of the basic weight is 1.0, and the value range of the classification confidence is [0, 1].
[0169] For example, if baseweight=1.0 and confidence=0.8, combined with Table 2, in a vertical composition, the semantic scene weight value is calculated as follows:
[0170] Character: Wscene = 1.0 + 0.25 × 0.8
[0171] Scenery: Wscene = 1.0-0.10×0.8
[0172] Document: Wscene=1.0+0.30×0.8
[0173] Food: Wscene = 1.0 + 0.15 × 0.8
[0174] Building: Wscene=1.0-0.05×0.8
[0175] In horizontal composition, the semantic scene weight value is calculated as follows:
[0176] Character: Wscene=1.0-0.15×0.8
[0177] Scenery: Wscene = 1.0 + 0.20 × 0.8
[0178] File: Wscene=1.0-0.25×0.8
[0179] Food: Wscene = 1.0 - 0.10 × 0.8
[0180] Building: Wscene = 1.0 + 0.18 × 0.8
[0181] Furthermore, semantic scene weights can also take user preferences into account. Different users develop specific compositional habits when shooting. For example, some prefer vertical compositions when capturing food or taking selfies, aligning with social media dissemination and handheld viewing. By analyzing and respecting these individual preferences, composition decisions can be more closely aligned with actual user needs, improving both the shooting experience and the visual quality.
[0182] S235 , performing feature fusion weighted summation on the target aspect ratio, edge density ratio, and semantic scene weight value of the preview frame to obtain a decision score.
[0183] The decision score can be calculated using the following formula:
[0184] Decision_Score=αWaspect+βWedge+γWscene
[0185] Where Decision_Score is the decision score, Waspect is the target aspect ratio, α is the weight of the target aspect ratio, Wedge is the edge density ratio, β is the weight of the edge density ratio, Wscene is the semantic scene weight value, and γ is the weight of the semantic scene weight value.
[0186] The weight of each feature is determined according to the importance of the feature, for example, α=0.4, β=0.35, and γ=0.25.
[0187] The shooting method provided in the embodiment of the present application uses a lightweight model to perform real-time analysis of preview frames in the preview stream, reflects the target aspect ratio of the preview frame composition decision, refines the edge density ratio of the preview frame composition decision, and adjusts the preview frame composition decision in combination with the scene category. The above features are integrated to obtain a decision score for indicating the composition direction (horizontal, vertical) suitable for the shooting target in the preview frame, thereby improving the generalization ability and robustness of the shooting method.
[0188] Figure 8 A process architecture diagram corresponding to a shooting method provided in an embodiment of the present application is shown. Figure 8 As shown in the figure, the process framework includes four parts: target aspect ratio calculation, edge density ratio calculation, semantic scene weight value mapping, and dynamic fusion and decision making.
[0189] The target aspect ratio calculation mainly includes aspect ratio feature extraction and multi-target fusion strategy. The first aspect ratio corresponding to the target is determined by photographing the outline of the target, and the first aspect ratios of one or more targets are integrated to obtain the target aspect ratio.
[0190] The edge density ratio calculation mainly includes real-time edge analysis, capturing target edges and edge density ratio calculation; statistical edge density ratio is used to supplement the characteristic dimension of the captured target.
[0191] In the semantic scene weight value mapping, it mainly includes the preset image classification model and dynamic weight adjustment, identifying scene categories, adapting the weights under different categories of scenes, and improving the accuracy of preview frame analysis.
[0192] In dynamic fusion and decision-making, it mainly includes weight fusion calculation and decision threshold analysis. After the decision score is obtained by fusing the above features, dual thresholds and adaptive adjustment of thresholds are used to analyze the decision score and adjust the shooting direction of the shooting target.
[0193] In some embodiments, when photographing a target, there may be a problem of incomplete target. In order to improve the integrity of the target in the shooting results during the above-mentioned horizontal and vertical switching process, the user can also be guided to adjust the shooting angle through cropping analysis of the target. Figure 9 A schematic diagram of a process flow of another shooting method provided in an embodiment of the present application is shown. Figure 9 As shown, for Figure 5 After the shooting target is determined in step 210, the following steps are also included:
[0194] S310: Analyze the preview frame based on the target detection algorithm to obtain the truncation degree of the captured target.
[0195] After analyzing the preview frames based on the target detection algorithm and generating the target bounding box for each target, the degree of truncation of the target can also be determined. The degree of truncation of the target refers to the area where the key areas of the target (such as the facial features of a person or the main body of a building) exceed the boundaries of the picture.
[0196] For example, the object detection algorithm is a modified YOLOv5s model. This model is lightweight and improved by adding an edge-aware convolution layer to improve the detection accuracy of truncated objects (such as partial people or buildings) at the image boundaries. The degree of truncation of the captured object can be determined by the intersection over union (IoU) between the object's bounding box and the image boundary. The truncation degree represents the proportion of the captured object that exceeds the preview frame boundary.
[0197] S320: If the truncation degree of the captured object is greater than the preset safety distance threshold, determine that the preview frame is in a truncation state.
[0198] The preset safety distance threshold can be a fixed threshold or can be dynamically adjusted according to the scene type. For example, for a human scene, in order to protect the integrity of the human face, the preset safety distance threshold can be reduced to, for example, 5%; for an architectural / landscape scene, some white space can be retained and the preset safety distance threshold can be increased to, for example, 10% to 15%.
[0199] The preset safety distance threshold can also be adaptively adjusted based on the size of the target and the movement speed collected by the sensor in the wearable device (for example, the movement speed caused by turning the head or turning around), reducing false triggering.
[0200] S330: When the preview frame is in a truncated state, generate a direction prompt voice, and output the direction prompt voice through a target speaker.
[0201] The directional voice prompt is used to indicate camera angle adjustment, that is, the directional voice prompt is used to remind the user to adjust the camera angle. If there is only a single target speaker, the directional voice prompt is output directly. If there are multiple target speakers, the directional voice prompt is output to one or more of the target speakers based on the angle adjustment corresponding to the directional voice prompt.
[0202] In some embodiments, the target speaker includes a first speaker and a second speaker, and the first speaker and the second speaker are respectively arranged at different positions of the wearable device, and the direction prompt voice is played through the target speaker, including: based on the direction information carried by the direction prompt voice; when the direction information is the first direction, determine the first speaker corresponding to the first direction, and play the direction prompt voice through the first speaker; when the direction information is the second direction, determine the second speaker corresponding to the second direction, and play the direction prompt voice through the second speaker.
[0203] For example, when the left boundary is cut off, the direction information is left. If the left side of the target exceeds the safe distance (for example, the left half of the face is out of the frame), a prompt "Please turn your head slightly to the right" is played in the user's right ear through the right speaker, simulating a sound source from the right, guiding the user to turn right naturally. When the right boundary is cut off, a direction prompt is played in the user's left ear through the left speaker, guiding the user to turn left.
[0204] In some embodiments, the target speaker is a bone conduction headset, which uses the Head-Related Transfer Function (HRTF) to convert the directional prompt voice into binaural sound intensity difference and time difference, enhance directional perception, and realize spatial audio coding. In addition, the bone conduction headset can also block ambient sound to ensure the safety of users when using it outdoors.
[0205] It should be understood that the above clipping analysis of the shooting target adopts a real-time update mechanism. For example, the clipping analysis of the shooting target in the preview frame is performed once every 200ms: the deviation between the current clipping state and the safe distance; the angular velocity collected by the IMU represents the rate of change of the user's head posture.
[0206] It can also use Kalman filtering to fuse data collected by the IMU to distinguish between the user's active head turning and unintentional shaking, thereby reducing misjudgment.
[0207] The embodiment of the present application analyzes the preview frame through the target detection algorithm, and when it is identified that there is a boundary truncation of the shooting target in the preview frame, generates a direction prompt voice for head posture correction, and improves the integrity of the shooting target through spatialized audio prompts.
[0208] The shooting method provided in the embodiment of the present application realizes millisecond-level horizontal / vertical adaptive switching through feature analysis of the preview frame by a lightweight model; combined with a sub-second closed-loop control mechanism, it improves the completeness of the shooting target, while reducing the user's learning time for operating the wearable device, improving the user experience, and forming a complete composition link of perception, decision-making and guidance.
[0209] It should be noted that in the embodiments of the present application, "greater than" can be replaced by "greater than or equal to", "less than or equal to" can be replaced by "less than", or "greater than or equal to" can be replaced by "greater than", and "less than" can be replaced by "less than or equal to".
[0210] It should be understood that the order of execution of the processes in the above embodiments does not necessarily imply a specific order of execution. The order of execution of the processes is determined by their functions and inherent logic, and does not constitute any limitation on the implementation of the embodiments of the present invention. The various embodiments described herein may be independent solutions or combined according to their inherent logic, and all such solutions fall within the scope of protection of this application.
[0211] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0212] The embodiment of the present application can divide the functional modules of the wearable device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other feasible division methods in actual implementation. The following is an example of dividing each functional module according to each function.
[0213] Figure 10 FIG1 shows a schematic diagram of a photographing device provided in an embodiment of the present application. The photographing device 400 can be used to perform the actions performed by the wearable device in the above method embodiment. Figure 10 As shown, the shooting device 400 includes a target determination unit 410, a preview unit 420, a multi-dimensional feature analysis unit 430 and a shooting direction analysis unit 440.
[0214] The target determination unit 410 is used to determine a shooting target.
[0215] The preview unit 420 is used to obtain a preview frame corresponding to the shooting target. The preview frame is a single frame of image data in the target preview stream captured by the camera, and the preview frame is not displayed during the capture process.
[0216] The multi-dimensional feature analysis unit 430 is configured to fuse preset visual features of the preview frame to obtain a decision score for the preview frame; the preset visual features include at least two of a target aspect ratio, an edge density ratio, and an image scene category, wherein the target aspect ratio represents the aspect ratio of the captured object in the preview frame, the edge density ratio represents the ratio of the total vertical edge intensity to the total horizontal edge intensity in the preview frame, and the image scene category represents the scene category of the content corresponding to the preview frame;
[0217] The shooting direction analysis unit 440 is used to analyze the decision score based on the dynamic threshold strategy and adjust the shooting direction of the shooting target. The dynamic threshold strategy dynamically adjusts the decision threshold according to the motion blur index.
[0218] The shooting device 400 provided in the embodiment of the present application is used to execute the shooting method of the above embodiment. The technical principles and technical effects are similar and will not be repeated here.
[0219] It should be noted that the shooting device 400 of the above application is embodied in the form of a functional unit. The term "unit" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.
[0220] The present application provides a computer program product that, when executed on a wearable device, causes the wearable device to execute the technical solution in the above embodiment. The implementation principle and technical effects are similar to those of the above method-related embodiments and will not be further described here.
[0221] The embodiment of the present application provides a readable storage medium, which contains instructions. When the instructions are executed on a wearable device, the wearable device executes the technical solution of the above embodiment. The implementation principle and technical effect are similar and will not be repeated here.
[0222] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0223] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0224] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is merely a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0225] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0226] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0227] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A shooting method, characterized in that: Applied to a wearable device, the wearable device includes a camera; the shooting method includes: Determine the shooting target; Acquire a preview frame corresponding to the shooting target, where the preview frame is a single frame of image data in the target preview stream captured by the camera, and the preview frame is not displayed during the capture process; fusing preset visual features of the preview frame to obtain a decision score for the preview frame; the preset visual features include at least two of a target aspect ratio, an edge density ratio, and an image scene category, wherein the target aspect ratio represents the aspect ratio of the photographed object in the preview frame, the edge density ratio represents the ratio of the total vertical edge intensity to the total horizontal edge intensity in the preview frame, and the image scene category represents the scene category of the content corresponding to the preview frame; The decision score is analyzed based on a dynamic threshold strategy to adjust the shooting direction of the shooting target. The dynamic threshold strategy dynamically adjusts the decision threshold according to the motion blur index.
2. The shooting method according to claim 1, wherein: The decision threshold includes a first threshold and a second threshold, and the first threshold is greater than the second threshold. The analyzing the decision score based on the dynamic threshold strategy and adjusting the shooting direction of the shooting target includes: When the decision scores corresponding to the preview frames of M consecutive frames are all greater than the first threshold, adjusting the shooting direction to a portrait mode; When the decision scores corresponding to the preview frames of M consecutive frames are all smaller than the second threshold, the shooting direction is adjusted to a horizontal mode; wherein M is a positive integer greater than 1.
3. The shooting method according to claim 2, characterized in that: The analyzing the decision score based on the dynamic threshold strategy and adjusting the shooting direction of the shooting target further includes: If the decision scores corresponding to the preview frames of the consecutive M frames meet the first condition or the second condition, maintaining the shooting direction; The first condition includes: the decision scores corresponding to the preview frames of the consecutive M frames are all within a numerical range between the first threshold and the second threshold; The second condition includes: the decision scores corresponding to the preview frames of the continuous M frames are not all greater than the first threshold, not all less than the second threshold, and not all within the numerical range between the first threshold and the second threshold.
4. The shooting method according to any one of claims 1 to 3, characterized in that: The preset visual features include the target aspect ratio, the edge density ratio, and the image scene category, and fusing the preset visual features of the preview frame to obtain a decision score for the preview frame, including: Analyzing the preview frame based on a segmentation algorithm to obtain a contour of the photographed target and determining the target aspect ratio of the preview frame; Analyzing the preview frame based on an edge detection algorithm and directional density to obtain the edge density ratio of the preview frame; Determining the image scene category corresponding to the preview frame based on a preset image classification model, and analyzing the image scene category, the preset mapping relationship between the preset scene category and the horizontal weight addition and the vertical weight addition, and the classification confidence to determine the semantic scene weight value of the preview frame; Performing feature fusion weighted summation on the target aspect ratio, the edge density ratio, and the semantic scene weight value of the preview frame to obtain the decision score.
5. The shooting method according to claim 4, characterized in that: Analyzing the preview frame based on a segmentation algorithm to obtain the outline of the photographed target and determining the target aspect ratio of the preview frame includes: Separating the outline of the shooting target from the preview frame based on a segmentation algorithm; Dividing the height of the outline by the width of the outline to obtain a first aspect ratio; When there is a single photographic target in the preview frame, determining the target aspect ratio of the preview frame based on the first aspect ratio of the photographic target; When there are multiple shooting targets in the preview frame, the target aspect ratio of the preview frame is determined based on the first aspect ratio of each shooting target and the target saliency weight corresponding to each shooting target. The target saliency weight is used to represent the quantified weight of the visual saliency of the shooting target in the preview frame, where the visual saliency includes at least size and position.
6. The shooting method according to claim 4, characterized in that: The analyzing the preview frame based on the edge detection algorithm and the direction density to obtain the edge density ratio of the preview frame includes: Determining the horizontal gradient and the vertical gradient of the preview frame based on a Sobel operator; Dividing the preview frame into a plurality of preset grids, and determining the total horizontal edge strength and the total vertical edge strength of the preset grids; determining a density ratio based on the total horizontal edge intensity and the total vertical edge intensity of each of the preset grids; Segment-wise conversion is performed based on the density ratio to obtain the edge density ratio of the preview frame.
7. The shooting method according to any one of claims 1 to 3, characterized in that: The wearable device further includes a microphone, and determining the shooting target includes: When the microphone captures a voice command, determining a user intent corresponding to the voice command; When the user intention is a shooting operation, parsing a target object from the voice instruction to obtain the shooting target; The initial preview stream collected by the camera is identified based on a target detection algorithm. After the shooting target is identified in the initial preview stream, the preview frame and the target preview stream corresponding to the shooting target are determined.
8. The shooting method according to claim 7, characterized in that: The wearable device further includes a target speaker, and after determining the shooting target, further includes: Analyzing the preview frame based on the target detection algorithm to obtain a truncation degree of the photographed target, wherein the truncation degree is used to represent a proportion of the photographed target exceeding a preview frame boundary; If the truncation degree of the photographed object is greater than a preset safety distance threshold, determining that the preview frame is in a truncation state; When the preview frame is in a truncated state, a direction prompt voice is generated and outputted through a target speaker, wherein the direction prompt voice is used to instruct adjustment of a shooting angle.
9. The shooting method according to claim 8, characterized in that: The target speaker includes a first speaker and a second speaker, the first speaker and the second speaker are respectively arranged at different positions of the wearable device, and the direction prompt voice is played through the target speaker, including: Based on the direction information carried by the direction prompt voice; When the direction information is a first direction, determining the first speaker corresponding to the first direction, and playing the direction prompt voice through the first speaker; When the direction information is a second direction, the second speaker corresponding to the second direction is determined, and the direction prompt voice is played through the second speaker.
10. A wearable device, characterized in that: The wearable device includes a camera and a processor, and the processor is configured to execute the shooting method according to any one of claims 1 to 9.
Citation Information
Cited By
Gimbal horizontal and vertical screen automatic switching method and device, gimbal and storage medium
CN122395450A