Earphone control method, earphone and storage medium

By integrating a video module into the headphones to detect facial occlusion in real time and output prompts, the problem of shooting experience caused by facial occlusion during headphone shooting is solved. This enables instant adjustments before shooting, avoids repeated shooting, and improves the user experience.

CN121985084APending Publication Date: 2026-05-05GEER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GEER TECH CO LTD
Filing Date
2026-01-05
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The headphones were only discovered to be obstructing the face after the shoot was completed, which affected the shooting experience.

Method used

By integrating a video module into the headphones, the system can detect facial occlusion areas in real time and output occlusion alerts, including voice or vibration alerts, when preset conditions are met.

Benefits of technology

Timely identification of facial occlusion during image acquisition avoids repeated shooting, optimizes user experience, and improves the smoothness and effectiveness of shooting operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985084A_ABST
    Figure CN121985084A_ABST
Patent Text Reader

Abstract

The invention discloses an earphone control method, an earphone and a storage medium, and relates to the technical field of communication, the earphone control method is applied to the earphone, the earphone comprises a video module, and the earphone control method comprises the following steps: in response to an image acquisition process for detecting a face shielding state, determining a face shielding area according to a video image acquired by the video module; and if the face shielding area satisfies a preset shielding condition, outputting shielding prompt information of the earphone. Through image acquisition and occlusion area judgment logic of the video module, linkage of occlusion state detection and response process triggering is realized, so that the earphone can output the occlusion prompt according to the actual face occlusion condition of the user, the condition that the occlusion is found after shooting is completed is avoided, and the shooting experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a control method for headphones, headphones, and a storage medium. Background Technology

[0002] In addition to playing music and making calls, headphones equipped with cameras can also take photos and videos, catering to diverse needs such as recording daily snippets, outdoor adventures, and real-time video sharing.

[0003] When wearing headphones, they need to fit snugly against the ear area, so it is inevitable that the user's face will be captured in the photo. At the same time, users usually only discover the problem of occlusion when reviewing the images or videos after taking the photo, which affects the shooting experience.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a control method for headphones, headphones, and storage medium, aiming to solve the technical problem that the obstruction is only discovered after the image is captured, which affects the shooting experience.

[0006] To achieve the above objectives, this application proposes a control method for headphones, which is applied to headphones equipped with a video module. The method includes: In response to the image acquisition process that detects facial occlusion, the facial occlusion area is determined based on the video image acquired by the video module. If the face occlusion area meets the preset occlusion conditions, the occlusion prompt information of the headphones is output.

[0007] In one embodiment, the step of determining the facial occlusion area based on the video image captured by the video module includes: Determine the detection direction corresponding to the headphone unit where the video module is located; If the video image has an occluded area, skin color region segmentation and connected component detection are performed on the video image based on the detection direction to obtain the facial occlusion area of ​​the video image.

[0008] In one embodiment, the step of performing skin color region segmentation and connected component detection on the video image based on the detection direction to obtain the facial occlusion region of the video image if an occlusion region exists in the video image includes: Extract the skin color mask from the video image; Using the detection direction as a reference, the pixels of the skin color mask are detected row by row to obtain the initial pixels; The initial pixel is used as the seed point to execute a region generation algorithm to obtain the facial occlusion region.

[0009] In one embodiment, prior to the step of determining the facial occlusion region based on the video image acquired by the video module in response to the image acquisition process of detecting facial occlusion, the control method of the headphones further includes: If the preset image acquisition conditions are met, the image acquisition process is triggered: The preset image acquisition conditions include at least the following: The earphone unit of the earphone was detected to be in a wearing state; Received voice command for image acquisition; Received the client's startup command; The image acquisition command sent by the client has been received.

[0010] In one embodiment, the step of outputting an occlusion warning message for the headphones if the face occlusion area meets a preset occlusion condition includes: If the face occlusion area meets the preset occlusion conditions, determine the occlusion position corresponding to the face occlusion area; Based on the earphone unit corresponding to the obstruction position, output obstruction prompt information and obstruction range adjustment information.

[0011] In one embodiment, before the step of outputting an occlusion prompt message from the earphone if the face occlusion area meets a preset occlusion condition, the earphone control method further includes: Obtain the percentage area of ​​the face occlusion region in the video image, and / or obtain the image overlap range between the earphone units of the headphones; If the area of ​​the face occlusion is greater than the preset area, and / or the area of ​​the face occlusion is greater than the area of ​​the image overlap, it is determined that the area of ​​the face occlusion meets the preset occlusion condition.

[0012] In one embodiment, after the step of determining the facial occlusion area based on the video image acquired by the video module in response to the image acquisition process of detecting facial occlusion, the control method of the headphones further includes: If the face occlusion area meets the preset occlusion conditions, the magnified image acquired by the video module based on the preset magnification factor is obtained; If the facial occlusion area in the magnified image meets the target occlusion condition, an occlusion prompt message is output for the headphones. If the facial occlusion area in the magnified image does not meet the target occlusion condition, the preset magnification factor is determined to be the current image acquisition factor of the headphones.

[0013] In one embodiment, the step of determining the facial occlusion region based on the video image acquired by the video module in response to the image acquisition process of detecting facial occlusion includes: In response to the image acquisition process that detects facial occlusion, the wearing history information of the headphones is obtained, and the wearing history information includes at least the occlusion probability and occlusion range of the headphone unit; The occlusion detection threshold of the headphone unit is updated based on the occlusion probability and the occlusion range. If the detected occlusion area in the video image is larger than the occlusion detection threshold, the facial occlusion area is determined based on skin color region segmentation and connected component detection.

[0014] In addition, to achieve the above objectives, this application also proposes an earphone, the earphone comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the earphone control method described above.

[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the headphone control method described above.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: When the earphone triggers the image acquisition process, it first determines the facial occlusion area based on the currently acquired image. If the facial occlusion area meets the occlusion conditions, it triggers the occlusion response process. Thus, in the real-time process of image acquisition, it synchronously identifies the facial occlusion area and determines whether the preset occlusion conditions are met, and outputs the occlusion prompt to the earphone in a timely manner. This allows users to know the facial occlusion situation before actual shooting and make timely adjustments, avoiding the need to reshoot due to invalid shooting content caused by occlusion. This optimizes the user's image shooting experience and improves the smoothness of shooting operation and the effectiveness of shooting results. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the first embodiment of the headphone control method of this application; Figure 2 This is a schematic diagram of the facial region features of the left ear image in the headphone control method of this application; Figure 3 This is a schematic diagram of the facial region features of the right ear image in the headphone control method of this application; Figure 4 This is a flowchart illustrating the second embodiment of the headphone control method of this application. Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the headphone control method in this application embodiment.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] The main solution of this application embodiment is: in response to the image acquisition process of detecting facial occlusion, the facial occlusion area is determined according to the video image acquired by the video module; If the face occlusion area meets the preset occlusion conditions, the occlusion prompt information of the headphones is output.

[0023] In this embodiment, for ease of description, the following description uses headphones as the execution subject.

[0024] Understandably, in traditional headphone shooting scenarios, either there is no occlusion detection function, requiring users to review the image after shooting to discover facial occlusion issues, resulting in invalid shots and a reduced experience, or it requires relying on external devices or manually activating the detection function, increasing operational steps and failing to meet the core requirement of portable headphone shooting. Based on this, this application provides a solution that deeply couples the occlusion detection logic with the built-in video module of the headset, requiring no additional hardware support. It also utilizes the headset's own interactive capabilities, such as voice and vibration trigger responses, to adapt to scenarios where users may be busy with their hands and unable to view the terminal screen in real time during headset shooting, achieving synchronization of acquisition and detection. Finally, when the occlusion area meets the conditions, a headset occlusion prompt is output, thus breaking the traditional process of shooting, post-verification, and reshooting. This forms a closed-loop occlusion processing solution adapted to portable headset shooting scenarios, simplifying the operation logic, enhancing the independence and practicality of the headset shooting function, and improving the user experience.

[0025] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or headset capable of performing the above functions. The headset is a wireless headset and is equipped with a video module, i.e., a camera, which is typically located on the left and right earpieces. The following description uses the headset as an example to illustrate this embodiment and the subsequent embodiments.

[0026] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0027] This application provides a method for controlling headphones, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the headphone control method of this application.

[0028] In this embodiment, the control method for the headphones includes steps S10 to S20: Step S10: In response to the image acquisition process of detecting facial occlusion, the facial occlusion area is determined based on the video image acquired by the video module.

[0029] In this embodiment, the image acquisition process for detecting facial occlusion is typically triggered before the user is ready to take a normal photo. For example, the user may be wearing headphones, the headphones may be connected to a client such as a mobile phone, or the user may have opened an application for managing headphone photography after the headphones are connected to the client. Therefore, the image acquisition process is triggered when the headphones meet preset image acquisition conditions. These preset conditions include detecting that the headphones are being worn, receiving a voice command for image acquisition, receiving a start command from the client, and receiving an image acquisition command sent by the client.

[0030] When shooting with headphones, because headphones are always worn close to one side of the face, the area of ​​facial occlusion in the captured image always starts from one side, such as... Figure 2 In the left ear image shown, the face is always depicted starting from the far right of the image. Figure 3 In the right ear image shown, the face is always depicted starting from the far left. Furthermore, the facial region is always a continuous area in the image, with clear boundaries between the edges and the background. Therefore, facial region recognition can utilize complex models or simple image processing algorithms. Thus, after triggering this process, the headphone's video module pre-captures a set of images and, based on these pre-captured images, determines whether facial occlusion exists and calculates the size of the occluded area if occlusion is present.

[0031] As an optional implementation, the headphones, in addition to the video module, also include capacitive sensors. After the headphones are powered on, the capacitive sensor readings for the left and right earpieces are recorded as follows: and ,when At that time, it was assumed that the left ear was in a wearing state. At that time, it was assumed that the right ear was being worn. The wearing threshold can be determined through actual product testing. When either the left or right earphone unit is worn, the ISP (Image Signal Processor) of that earphone unit is actively activated to capture real-time images. If both ears are worn simultaneously, the ISPs of both earphones are activated simultaneously to capture real-time images. After capturing the real-time image, a simple model is first used to identify whether occlusion exists. If so, a pre-trained model is used for processing, such as a pre-trained YOLOv5 face detection and occlusion segmentation model, to identify whether the occluded area in the image is a facial occlusion area, and the specific location and size of the facial occlusion area. Understandably, since this image is only used for facial region detection, the image acquisition resolution can be set, such as setting the video module to capture lower pixel images, such as 640P or 480P images, to reduce power consumption.

[0032] As another alternative implementation, besides using deep learning models for recognition, facial occlusion detection can also be performed based on computer vision recognition. Specifically, after the headphones respond to the image acquisition process triggered by the user's voice, they control the video module to acquire the currently captured image. This image is then input into a Haar-like feature (rectangular features for object detection) face detector trained using the Adaboost algorithm (adaptive boosting based on ensemble learning). The detector locates the facial region and outputs key facial feature points (coordinates of the eyes, nose, and mouth). Next, a pre-stored complete facial template is called, and the detected facial region is normalized and matched with the complete facial template. The grayscale difference matrix between the two is calculated, a grayscale difference threshold is set, and regions with grayscale differences greater than the threshold are identified as occluded regions. Simultaneously, the coordinates and range of the occluded region are output.

[0033] Step S20: If the face occlusion area meets the preset occlusion conditions, output the occlusion prompt information to the headphones.

[0034] In this embodiment, the preset occlusion condition is a pre-defined quantitative condition used to determine whether a facial occlusion area in an image constitutes effective occlusion. The facial occlusion response process is used to prompt the user that, at the current headphone wearing angle, there is significant occlusion that prevents the images captured by both headphones from being effectively stitched together, or that the image captured by one headphone is severely occluded, resulting in an image that does not meet the user's shooting requirements. This response process can be triggered by voice prompts or vibration prompts.

[0035] As an optional implementation, the preset occlusion condition is that the proportion of the face occlusion area in the video image is greater than a preset area, and the face occlusion response process is a voice prompt. Before step S20, the proportion of the face occlusion area in the video image can be obtained first. If the proportion is greater than the preset area, it is determined that the face occlusion area meets the preset occlusion condition. After meeting the preset occlusion condition, the occlusion position corresponding to the face occlusion area is determined first, and according to the earphone unit corresponding to the occlusion position, occlusion prompt information and an occlusion diagram corresponding to the video image are output so that the user can adjust the earphone wearing position. For example, when the earphone is worn only in the left ear, if the proportion of the face occlusion area in the video image captured by the left earphone unit is greater than the preset 30%, it indicates that the left ear occlusion exceeds the threshold, and a prompt tone of "left ear is occluded" is played, while an occlusion diagram is output to the client. Similarly, when the earphone is worn in the right ear and the occlusion condition is met, a prompt tone of "right ear is occluded" is output. Furthermore, when both ears are worn simultaneously, a notification can be triggered if the facial occlusion area of ​​the image in either the left or right ear reaches a threshold, indicating that one earbud is occluded or both are occluded.

[0036] In another alternative implementation, due to the physical limitations of the earphone wearing area, the video modules of the two earphones are typically integrated on the outside of the ear stem or close to the auricle. The shooting angle is determined by the spatial distance between the two ears and the wearing angle. When taking photos simultaneously, the shooting fields of view of the two modules are divided into overlapping and non-overlapping sides due to the obstruction of the human head and spatial overlap. The non-overlapping side refers to the area covered by only one earphone, such as the edges of both sides of the head or the outside of the auricle, while the overlapping side is the intersection of the fields of view of the two earphone modules, which corresponds precisely to the core area of ​​the human face.

[0037] When two headphones simultaneously capture images, fusion is typically performed based on the overlapping area after capture. However, if facial occlusion occurs, and this occlusion area occupies a significant portion of the overlapping side, it prevents effective fusion of the two images. Furthermore, the fused image not only exhibits noticeable incompleteness, artifacts, or feature loss in the core facial area, resulting in an unclear and complete portrait image, but the feature loss caused by occlusion also disrupts the image fusion registration logic, preventing the fusion algorithm from achieving pixel-level precision alignment. Consequently, the final output image fails to meet the expected shooting effect and forces users to retake shots due to invalid results, negatively impacting the user experience in portable headphone shooting scenarios.

[0038] Therefore, before step S20, in addition to determining whether the face occlusion area meets the preset occlusion conditions by the proportion of the face occlusion area in the video image, it is also possible to determine whether the size of the face occlusion area is greater than the overlap range of the images captured by the left and right earphone units. Specifically, the image overlap range between the earphone units can be obtained, and then, if the face occlusion area is greater than the image overlap range, it is determined that the face occlusion area meets the preset occlusion conditions. The size of the image overlap range is usually greater than or equal to a preset area; for example, when the preset area accounts for 30%, the image overlap range is usually greater than or equal to 30%. When the face occlusion area is greater than the image overlap range, the left and right images cannot be effectively stitched together because there is no overlapping area, therefore, an occlusion prompt needs to be output.

[0039] Furthermore, in addition to judging by the proportion area and the image overlap range respectively, it is also possible to judge whether the proportion area is greater than the preset area and whether the face occlusion area is greater than the image overlap range. When both conditions are met, it is judged that the face occlusion area meets the preset occlusion conditions.

[0040] Optionally, in addition to outputting occlusion prompts, the system can also calculate the current occlusion ratio and output occlusion range adjustment information based on the occlusion ratio, such as informing the user that the occlusion range needs to be adjusted by "rotating 5° to the left, please confirm".

[0041] Understandably, when the face occlusion area does not meet the preset occlusion requirements, the headphones will not perform the corresponding prompt action and will capture an image of the current environment in response to subsequent shooting instructions.

[0042] This embodiment provides a headphone control method that integrates occlusion detection into the headphone system. By combining images captured by the headphone itself with facial features captured while the headphone is worn, it detects facial occlusion in real time. Based on a user-defined threshold, it alerts the user via audible prompts, preventing unusable photos taken while using the headphone from being obstructed. This achieves automatic occlusion detection and immediate response in headphone shooting scenarios. No additional user intervention is required, integrating the headphone and occlusion detection functions. This enables headphone with a camera to have a self-calibration function for shooting status, filling a technological gap in automatic occlusion control in headphone shooting scenarios. It not only improves the integrity of captured content and avoids invalid shots due to occlusion, but also eliminates the need to interrupt shooting for occlusion checks, enhancing the user shooting experience.

[0043] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 Step S10 also includes steps S11 to S12: Step S11: Determine the detection direction corresponding to the headphone unit where the video module is located.

[0044] In this embodiment, facial occlusion areas can be identified using skin color region segmentation combined with connected component detection. Connected component detection requires determining the image scanning direction to reduce computational load. Specifically, since the occluded facial area is on the right side of the image when the left earphone unit is capturing the image, scanning starts from the right, reducing pixel scanning of irrelevant areas and improving processing efficiency. Therefore, the detection direction can be determined using the earphone unit where the video module is located: the detection direction for the left earphone unit is right, and the detection direction for the right earphone unit is left.

[0045] Step S12: If there is an occluded area in the video image, perform skin color region segmentation and connected component detection on the video image based on the detection direction to obtain the facial occlusion area of ​​the video image.

[0046] In this embodiment, after identifying the occluded area in the image, the skin color mask of the video image can be extracted first. Then, based on the detection direction, the pixels of the skin color mask are detected row by row to obtain the initial pixels. Then, the region generation algorithm is executed with the initial pixels as the seed point area to obtain the facial occlusion area.

[0047] Specifically, when extracting the skin tone mask, the original video image can be converted to... The space is then used to extract the skin color mask based on the following calculation formula: ,in, It refers to the range of skin color in the acquired image samples. Represents the luminance component. Indicates the red difference component. This represents the blue difference component.

[0048] After extracting the skin color mask, for the left ear image, the detection direction is from right to left. Therefore, the scan starts from the right boundary and proceeds to the left to obtain the results of the skin color mask. Detect the first pixel line by line .by Execute a region generation algorithm for seed points to mark contiguous regions. This refers to the area where the face is covered.

[0049] When calculating the area of ​​this region, it is necessary to obtain the bounding rectangle of the facial region. Then the area of ​​the facial region is obtained as follows: .

[0050] Finally, the proportion of the facial region in the image is calculated. ,in This represents the total area of ​​the image on the left.

[0051] This embodiment provides a method for controlling headphones. By determining the corresponding detection direction of the headphone unit corresponding to the video module, the method can match the shooting perspective difference between the left and right headphones, making skin color region segmentation more accurate and focusing on the face region under that perspective. At the same time, the method simplifies the scanning and calculation logic of connected region detection by using directional detection, adapts to the limited computing power resources of the headphones, provides reliable occlusion information for subsequent dual-unit image fusion, reduces the fusion failure problem caused by occlusion recognition deviation, and optimizes the headphone shooting experience.

[0052] Based on the first embodiment of this application, in the third embodiment of this application, the same or similar content as the first embodiment can be referred to the above description, and will not be repeated hereafter. Furthermore, when shooting with both earphones simultaneously, it is usually necessary to stitch and merge the captured images. However, when shooting with only one earphone, the current shooting magnification can be automatically increased, thereby cropping and focusing the core area of ​​the image while removing the edge obstructions to the face when the earphone is worn.

[0053] Therefore, in this embodiment, when the facial occlusion area meets the preset occlusion conditions in normal shooting mode, the earphone's video module can also acquire a magnified image based on a preset magnification factor. By reducing the image acquisition range of a single ear, the proportion of the current occlusion area relative to the entire image is reduced, so that the earphone can directly acquire images based on this magnification factor in subsequent shots, reducing prompts and fundamentally avoiding situations where occlusion is discovered after shooting and forced reshooting. To improve image clarity, this magnification factor is typically between 1.5 and 3 times, but can also be adjusted based on actual needs.

[0054] Therefore, after step S10, if the face occlusion area meets the preset occlusion conditions, such as having a proportion area greater than a preset area, and / or the face occlusion area being larger than the image overlap range, then it is necessary to acquire a magnified image collected by the video module based on a preset magnification factor. Subsequently, it is determined whether the face occlusion area in the magnified image meets the target occlusion conditions, i.e., whether the face occlusion area in the magnified image also meets the above conditions. If so, the face occlusion process is triggered.

[0055] Furthermore, if the occlusion range in normal mode exceeds the preset occlusion condition, but the occlusion range after magnification based on the preset magnification does not exceed the condition, the system determines the preset magnification to be the current image acquisition magnification of the headphones. This avoids triggering unnecessary responses due to occlusion in normal shooting mode, reduces interference with the user's shooting operation, and eliminates the need for the user to frequently adjust the headphone position, improving the practicality and flexibility of shooting and better meeting the needs of shooting with headphones.

[0056] This embodiment provides a control method for headphones. When no image stitching or fusion is required after single-ear or dual-ear shooting, if the detection finds that the occlusion area in normal shooting mode is greater than a preset threshold, an image is captured based on a preset magnification. The method then determines whether the magnified image meets similar occlusion conditions. If so, a face occlusion response process is directly triggered; otherwise, subsequent image capture is performed based on the magnification. This reduces the impact of the face occlusion area by adjusting the image magnification when occlusion occurs in normal shooting mode, rather than directly issuing a warning. This allows users to complete effective shooting without manually adjusting the headphone position, fundamentally avoiding the problem of discovering occlusion only after shooting and being forced to reshoot. It saves shooting time, simplifies operation steps, and ensures that the final image meets usage requirements, improving the user experience and practicality of portable headphone shooting.

[0057] Based on the first embodiment of this application, in the fourth embodiment of this application, the content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. In addition, step S10 further includes steps S13 to S15: Step S13: In response to the image acquisition process of detecting facial occlusion, obtain the wearing history information of the headphones.

[0058] In this embodiment, the wearing history information includes at least the occlusion probability and occlusion range of the earphone units. Specifically, after at least one earphone unit detects that it is being worn, a wearing history information acquisition command is triggered, and wearing history information is retrieved from the system's non-volatile storage module. This includes the occlusion probability and corresponding occlusion range of the left and right earphone units each time they are worn. Simultaneously, the historical data is validated to remove abnormal data, such as invalid records from failed acquisitions. By learning the occlusion probability and occlusion range of the earphone units during the user's past wear, the system learns the user's wearing habits. This allows for the optimization of the occlusion detection threshold for each earphone unit based on these habits, improving the accuracy and personalization of occlusion detection, and reducing false positives and invalid prompts.

[0059] Step S14: Update the occlusion detection threshold of the headphone unit based on the occlusion probability and occlusion range.

[0060] In this embodiment, the occlusion detection threshold is used to initially determine whether occlusion exists. For example, the occlusion detection threshold is set to 10% of the image area. If the occlusion area does not exceed this threshold, it can be considered that there is no facial occlusion area. The occlusion detection threshold can be adjusted using a weighted fusion method, based on the weighted proportions and scores of the occlusion probability and the occlusion range. This allows for personalized dynamic calibration of the occlusion detection threshold, making the threshold adjustment more aligned with individual user wearing habits.

[0061] Step S15: If the occlusion area detected in the video image is greater than the occlusion detection threshold, the current occlusion area is determined to be a facial occlusion area.

[0062] After obtaining the video image, it can be determined whether the proportion of the occluded area in the image is greater than the occlusion detection threshold, thereby determining whether a facial occlusion area exists. If so, the facial occlusion area is determined based on skin color region segmentation and connected component detection; otherwise, the facial occlusion area is considered not to exist.

[0063] This embodiment provides a control method for headphones. The occlusion detection threshold is updated based on wearing history information such as the occlusion probability and occlusion range of the headphone unit. This allows the occlusion detection threshold to dynamically adapt to the individual wearing habits of the user, avoiding misjudgment or missed detection caused by a fixed general threshold, improving the recognition accuracy of facial occlusion areas, and reducing unnecessary occlusion prompts from interfering with the shooting operation.

[0064] This application provides an earphone, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the earphone control method described in the first embodiment above.

[0065] The following is for reference. Figure 5 The diagram shows a structural schematic of an earphone suitable for implementing embodiments of this application. Figure 5 The headphones shown are merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0066] like Figure 5 As shown, the headphones may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for headphone operation. The processing device 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the headset to communicate wirelessly or wiredly with other devices to exchange data. Although headsets with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0067] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0068] The earphone provided in this application, using the earphone control method in the above embodiments, can solve the technical problem of discovering obstruction only after image capture is completed, thus affecting the shooting experience. Compared with the prior art, the beneficial effects of the earphone provided in this application are the same as those of the earphone control method provided in the above embodiments, and other technical features of the earphone are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0069] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0070] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0071] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the headphone control method in the above embodiments.

[0072] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM, or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0073] The aforementioned computer-readable storage medium may be included in the headphones; or it may exist independently and not assembled into the headphones.

[0074] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the headphones, cause the headphones to: In response to the image acquisition process that detects facial occlusion, the facial occlusion area is determined based on the video image acquired by the video module. If the face occlusion area meets the preset occlusion conditions, the occlusion prompt information of the headphones is output.

[0075] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0077] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0078] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described headphone control method. This solves the technical problem of discovering obstruction only after image capture is complete, thus affecting the shooting experience. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the headphone control method provided in the above embodiments, and will not be repeated here.

[0079] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for controlling headphones, characterized in that, Applied to headphones, the headphones are equipped with a video module, and the control method of the headphones includes: In response to the image acquisition process that detects facial occlusion, the facial occlusion area is determined based on the video image acquired by the video module. If the face occlusion area meets the preset occlusion conditions, the occlusion prompt information of the headphones is output.

2. The headphone control method as described in claim 1, characterized in that, The step of determining the facial occlusion area based on the video image captured by the video module includes: Determine the detection direction corresponding to the headphone unit where the video module is located; If the video image has an occluded area, skin color region segmentation and connected component detection are performed on the video image based on the detection direction to obtain the facial occlusion area of ​​the video image.

3. The headphone control method as described in claim 2, characterized in that, The step of performing skin color region segmentation and connected component detection on the video image based on the detection direction to obtain the facial occlusion region of the video image if there is an occlusion region in the video image includes: Extract the skin color mask from the video image; Using the detection direction as a reference, the pixels of the skin color mask are detected row by row to obtain the initial pixels; The initial pixel is used as the seed point to execute a region generation algorithm to obtain the facial occlusion region.

4. The headphone control method as described in claim 1, characterized in that, Before the step of determining the facial occlusion area based on the video image acquired by the video module in response to the image acquisition process for detecting facial occlusion, the control method for the headphones further includes: If the preset image acquisition conditions are met, the image acquisition process is triggered: The preset image acquisition conditions include at least the following: The earphone unit of the earphone was detected to be in a wearing state; Received voice command for image acquisition; Received the client's startup command; The image acquisition command sent by the client has been received.

5. The headphone control method as described in claim 1, characterized in that, The step of outputting an occlusion warning message for the headphones if the face occlusion area meets the preset occlusion conditions includes: If the face occlusion area meets the preset occlusion conditions, determine the occlusion position corresponding to the face occlusion area; Based on the earphone unit corresponding to the obstruction position, output obstruction prompt information and obstruction range adjustment information.

6. The control method for headphones as described in any one of claims 1 to 5, characterized in that, Before the step of outputting an occlusion prompt message to the headphones if the face occlusion area meets the preset occlusion conditions, the headphone control method further includes: Obtain the percentage area of ​​the face occlusion region in the video image, and / or obtain the image overlap range between the earphone units of the headphones; If the area of ​​the face occlusion is greater than the preset area, and / or the area of ​​the face occlusion is greater than the area of ​​the image overlap, it is determined that the area of ​​the face occlusion meets the preset occlusion condition.

7. The headphone control method as described in claim 1, characterized in that, After the step of determining the facial occlusion area based on the video image acquired by the video module in response to the image acquisition process for detecting facial occlusion, the control method for the headphones further includes: If the face occlusion area meets the preset occlusion conditions, the magnified image acquired by the video module based on the preset magnification factor is obtained; If the facial occlusion area in the magnified image meets the target occlusion condition, an occlusion prompt message is output for the headphones. If the facial occlusion area in the magnified image does not meet the target occlusion condition, the preset magnification factor is determined to be the current image acquisition factor of the headphones.

8. The headphone control method as described in claim 1, characterized in that, The step of determining the facial occlusion region based on the video image acquired by the video module in response to the image acquisition process for detecting facial occlusion includes: In response to the image acquisition process that detects facial occlusion, the wearing history information of the headphones is obtained, and the wearing history information includes at least the occlusion probability and occlusion range of the headphone unit; The occlusion detection threshold of the headphone unit is updated based on the occlusion probability and the occlusion range. If the detected occlusion area in the video image is larger than the occlusion detection threshold, the facial occlusion area is determined based on skin color region segmentation and connected component detection.

9. An earphone, characterized in that, The earphone includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the earphone control method as claimed in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the headphone control method as described in any one of claims 1 to 8.