Earphone control method, earphone and storage medium

The earbud control method optimizes image capture by dynamically switching between single-ear and dual-ear modes based on scene complexity, addressing power consumption issues and maintaining flexibility in low-complexity scenarios.

CN120321544APending Publication Date: 2025-07-15WEIFANG GOERTEK ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510412283.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In shooting scenes with low complexity static portraits or single subjects, the dual-ear shooting method increases unnecessary power consumption, resulting in reduced headphone battery life, frequent charging or interruption of use, reducing shooting flexibility.

Method used

The environmental image is collected through the first earphone unit, the scene to be photographed is recognized, and the earphones are controlled to perform single-ear or double-ear shooting actions according to the recognition results, and the shooting mode is optimized to improve flexibility.

Benefits of technology

It improves the shooting flexibility of the headphones in different scenarios, avoids the problems of excessive power consumption or poor shooting results, and improves battery life and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321544A_ABST
    Figure CN120321544A_ABST
Patent Text Reader

Abstract

The invention discloses an earphone control method, an earphone and a storage medium, and relates to the technical field of wireless earphones, the earphone control method is applied to the earphone, the earphone comprises a first earphone unit and a second earphone unit which are provided with video modules, and the earphone control method comprises the following steps: responding to a pre-shooting instruction; acquiring an environment image acquired by the first earphone unit; determining a to-be-shot scene according to an image recognition result of the environment image; and controlling the earphone to execute a single-ear shooting action or a double-ear shooting action according to a shooting demand corresponding to the scene to be shot. Therefore, in the shooting process, the to-be-shot scene and the shooting requirement of the to-be-shot scene are obtained based on the pre-shot environment image, and the earphone is controlled to execute the corresponding shooting action through the actual shooting requirement of the scene, so that the shooting flexibility is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of wireless earphones, and particularly to a control method for earphones, earphones and a storage medium. Background Art

[0002] Integrating a camera module on the earphone enables the earphone to support scenarios such as intelligent navigation, real-time translation, and AR interaction. As users' demands for the comprehensiveness and accuracy of image acquisition increase day by day, when taking pictures or videos based on the earphone, two earphones need to cooperate for shooting and processing.

[0003] However, in shooting scenarios with low complexity such as static portraits and single subjects, the dual-ear shooting method will increase unnecessary power consumption, reduce the battery life of the earphone, cause the earphone to need to be charged frequently or experience usage interruptions, and reduce the flexibility of shooting.

[0004] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a control method for earphones, earphones and a storage medium, aiming to solve the technical problem of poor flexibility in earphone shooting.

[0006] To achieve the above purpose, this application proposes a control method for earphones, which is applied to earphones. The earphones include a first earphone unit and a second earphone unit provided with a video module. The control method for the earphones includes:

[0007] Respond to a pre-shooting instruction and obtain the environmental image collected by the first earphone unit;

[0008] Determine the scene to be shot according to the image recognition result of the environmental image;

[0009] Control the earphone to perform a single-ear shooting action or a dual-ear shooting action according to the shooting requirements corresponding to the scene to be shot.

[0010] In one embodiment, the step of controlling the earphone to perform a single-ear shooting action or a dual-ear shooting action according to the shooting requirements corresponding to the scene to be shot includes:

[0011] If the shooting requirement is dual-ear shooting, obtain the wearing information of the first earphone unit and the second earphone unit;

[0012] If both the first earphone unit and the second earphone unit are in a worn state, control the earphone to perform a dual-ear shooting action;

[0013] Otherwise, control the first earphone unit to perform a single-ear shooting action.

[0014] In one embodiment, the step of controlling the earphone to perform a binaural shooting action includes:

[0015] Based on the first earphone unit, synchronize the image acquisition instruction to the second earphone unit. After receiving the image acquisition instruction, the second earphone unit activates the video module;

[0016] Control the video modules of the first earphone unit and the second earphone unit to perform image acquisition actions.

[0017] In one embodiment, after the step of controlling the earphone to perform a binaural shooting action, it further includes:

[0018] Obtain a first image and a second image, where the first image and the second image are images collected by the first earphone unit and the second earphone unit respectively;

[0019] Determine the overlapping area of the first image and the second image, and perform feature matching according to the feature information of the overlapping area to obtain a similarity;

[0020] If the similarity meets a preset similarity, synthesize the first image and the second image to obtain a target image corresponding to the binaural shooting action;

[0021] Otherwise, generate a prompt message for the wearing position adjustment of the earphone unit and / or inconsistent shooting.

[0022] In one embodiment, before the step of determining the scene to be shot according to the image recognition result of the environmental image, it further includes:

[0023] Based on a pre-trained model, identify and process the object information of the environmental image to obtain the number of objects in the environmental image;

[0024] Generate an image recognition result including the number of objects;

[0025] The step of determining the scene to be shot according to the image recognition result of the environmental image includes:

[0026] If the number of objects is greater than a preset number, determine that the scene to be shot is a first scene, and the shooting requirement of the first scene is binaural shooting;

[0027] Otherwise, determine that the scene to be shot is a second scene, and the shooting requirement of the second scene is monaural shooting.

[0028] In one embodiment, the step of determining the scene to be shot according to the image recognition result of the environmental image includes:

[0029] Obtain the voice information collected by the earphone;

[0030] Determine the scene to be photographed according to the image recognition result and the voice information.

[0031] In one embodiment, the step of determining the scene to be photographed according to the environmental image recognition result further includes:

[0032] Obtain the voice information and motion data collected by the earphone;

[0033] Determine the scene to be photographed according to the image recognition result, the voice information and the motion data.

[0034] In one embodiment, the shooting requirements include dual-ear shooting and single-ear shooting, and the shooting requirements are determined based on any one of a preset instruction, the object quantity information of the image recognition result, or user behavior habit self-learning.

[0035] In addition, to achieve the above object, the present application also proposes an earphone, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the control method of the earphone as described above.

[0036] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the control method of the earphone as described above.

[0037] One or more technical solutions proposed by the present application have at least the following technical effects:

[0038] During the shooting process, first pre-collect the image through the first earphone unit, then recognize the pre-collected image, and determine the scene that the user actually needs to shoot based on the recognition result. Finally, according to the shooting requirements corresponding to the scene to be photographed, control the earphone to perform single-ear shooting or dual-ear shooting. Based on this, it is possible to control the earphone to perform dual-ear shooting or single-ear shooting according to the actual shooting requirements of the scene, improving the flexibility of earphone shooting. Description of the Drawings

[0039] The accompanying drawings here are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0041] Figure 1 The flowchart provided for the first embodiment of the control method of the earphone of the present application;

[0042] Figure 2 The flowchart provided for the second embodiment of the control method of the earphone of the present application;

[0043] Figure 3 The optional flowchart of the control method of the earphone provided for the third embodiment of the present application;

[0044] Figure 4 The schematic diagram of the device structure of the hardware operating environment involved in the control method of the earphone in the embodiments of the present application.

[0045] The realization of the purpose of the present application, functional features and advantages will be further described in combination with the embodiments with reference to the accompanying drawings. Specific embodiments

[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0047] In this embodiment, for the convenience of description, the following will be described with the earphone as the execution subject.

[0048] Integrate a camera module on the earphone, so that the earphone supports scenarios such as intelligent navigation, real-time translation, and AR interaction. With the increasing demand of users for the comprehensiveness and accuracy of image acquisition, when taking pictures or videos based on the earphone, two earphones need to cooperate for shooting and processing.

[0049] However, in low-complexity shooting scenarios such as static portraits and single subjects, the binocular shooting method will increase unnecessary power consumption, reduce the battery life of the earphone, cause the earphone to need to be charged frequently or the use to be interrupted, and reduce the flexibility of shooting.

[0050] Based on this, the main solution of the embodiments of the present application is: in response to a pre-shooting instruction, obtain the environmental image collected by the first earphone unit;

[0051] Determine the scene to be shot according to the image recognition result of the environmental image;

[0052] Control the earphone to perform a single-ear shooting action or a binocular shooting action according to the shooting requirements corresponding to the scene to be shot.

[0053] Specifically, during the shooting process, the first earphone unit is first used to pre-capture an image. Subsequently, the pre-captured image is recognized, and the scene that the user actually needs to shoot is determined based on the recognition result. Finally, according to the shooting requirements corresponding to the scene to be shot, the earphone is controlled to perform single-ear shooting or double-ear shooting, so as to shoot based on the actual needs of the user and improve the shooting flexibility.

[0054] To better understand the technical solution of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific embodiments.

[0055] The embodiment of the present application provides a control method for an earphone, which is applied to an earphone. The earphone includes a first earphone unit and a second earphone unit provided with a video module, that is, both earphones can perform shooting. The earphone is a wireless earphone or a head-mounted earphone, and the video module is a micro camera module integrated in the earphone cavity. Please refer to Figure 1 , Figure 1 which is a schematic flow chart of the first embodiment of the control method for the earphone of the present application.

[0056] In this embodiment, the control method for the earphone includes steps S10 to S30:

[0057] Step S10, in response to a pre-shooting instruction, obtain an environmental image collected by the video module of the first earphone unit.

[0058] The pre-shooting instruction refers to a control instruction received through a specified key, a trigger signal received by a wireless transmission protocol, gesture information, or voice keyword recognition. The control instruction is triggered by the user clicking the shooting button on the earphone, clicking the pre-shooting function on the application connected to the earphone, and triggering by voice keywords such as performing a shooting check, etc. It should be noted that during the pre-shooting process, the first earphone unit is in a worn state.

[0059] In this embodiment, when the user clicks pre-shooting on the application on the mobile phone connected to the earphone, the main control chip of the earphone receives the pre-shooting instruction sent by the mobile phone through the Bluetooth protocol stack, and then controls the video module of the first earphone unit to perform a shooting action to collect the current environmental image. Among them, the first earphone unit can be the left earphone unit or the right earphone unit. At the same time, parameters such as the focusing parameter and resolution when the video module collects the environmental image can be pre-set or selected based on the current battery power. The present application does not make any limitations here.

[0060] Optionally, after the voice recognition module of the earphone recognizes keywords such as "start shooting" and "take a photo", it performs a shooting action through the video module of the first earphone unit.

[0061] When shooting based on the earphone, pre-shooting processing is performed through the first earphone unit to identify the pre-shot image, and then based on the recognition results such as environmental complexity (number of objects, color distribution) and the current location area, etc., the shooting requirements in the current scene are determined to improve shooting flexibility.

[0062] Step S20, determine the scene to be shot according to the image recognition result of the environmental image.

[0063] In this embodiment, after obtaining the environmental image, it is also necessary to perform recognition processing on the environmental image to determine the shooting requirements of the current scene to be shot based on the recognition result. The image recognition result at least includes the number of objects, color complexity, and the corresponding location / area of the environment in the environmental image, etc.

[0064] As an optional implementation manner of image recognition, the environmental image can be input into a pre-trained model, and image recognition processing is performed based on the pre-trained model to identify information such as things and people in the image. Subsequently, a judgment is made according to the recognized quantity and a preset quantity threshold to obtain the current scene to be shot and the complexity of the scene to be shot. That is, before determining the scene to be shot, the object information of the environmental image is recognized based on the pre-trained model to obtain the number of objects in the environmental image, and then an image recognition result including the number of objects is generated.

[0065] Furthermore, during the process of determining the scene to be shot, if the number of objects is greater than the preset number, the scene to be shot is determined as the first scene; otherwise, the scene to be shot is determined as the second scene. Among them, the shooting requirement of the first scene is binocular shooting, and the shooting requirement of the second scene is monaural shooting. It can be understood that in the scene to be shot, if the number of objects to be shot is small, it indicates that the current shooting scene is relatively simple. For example, in a simple scene with a solid color background, a minimalist stage, or a narrow passage including a long corridor, alley, or narrow street with a strong sense of visual extension but single elements, an image can be shot using one earphone without binocular shooting, reducing the headphone power consumption and improving shooting flexibility. In a complex scene, such as when the number of objects to be shot is large, it indicates that the current shooting scene is relatively complex. At this time, binocular shooting is required to improve the clarity and shooting range of the shot image and the shooting effect in a complex scene. Therefore, the complexity of the scene to be shot is determined through the image recognition result, so as to determine the actual shooting requirements based on the complexity, avoiding the situation of high power consumption caused by binocular shooting in a simple scene or poor shooting effect caused by monaural shooting in a complex scene, and improving the shooting flexibility of the earphone.

[0066] Optionally, during the image recognition process, a scene classification model constructed by a convolutional neural network can also be used to analyze and process the environmental image, so as to obtain the scene to be photographed. For example, the input layer of the model is a 224×224 normalized RGB image, and the output layer is a 6-dimensional probability vector, including scenes such as corridor / meeting room / outdoor / sports / low light / others. Based on this, after inputting the environmental image into the model, the scene type of the scene to be photographed currently located can be determined based on the probability results of each output layer. By determining the scene type of the scene to be photographed, the earphone can be controlled to execute corresponding photographing actions based on the photographing requirements associated with the scene type, improving the flexibility of photographing.

[0067] Optionally, the position in the current environment can also be determined through the image recognition result. For example, if the image recognition result is the famous welcoming pine tree in a certain scenic area, it can be determined that the current scene to be photographed is a tourism scene at this time. In addition, when performing pre-photographing, the real-time position information of the earphone can also be obtained, and the current scene to be photographed can be determined based on the image recognition result and the position information. For example, when the user is in a certain science and technology park and the image recognition result is a display screen, it can be considered that the scene to be photographed is a meeting scene at this time. Further, the complexity of the scene to be photographed can also be determined through the types and quantities of colors in the image recognition result, etc.

[0068] Step S30, according to the photographing requirements corresponding to the scene to be photographed, control the earphone to execute a single-ear photographing action or a double-ear photographing action.

[0069] The photographing requirements include double-ear photographing and single-ear photographing. The scene to be photographed is usually associated with photographing requirements, and the associated photographing requirements can be determined based on the environmental complexity. For example, simple scenes and complex scenes are respectively associated with the photographing requirements of single-ear photographing and double-ear photographing, that is, the photographing requirements are determined based on the object quantity information of the image recognition result.

[0070] Optionally, the photographing requirements can be determined according to a preset instruction. For example, in the meeting scene, sports scene, and outdoor scene set by the user, the photographing requirement is double-ear photographing, while in the corridor, low light and other scenes, the photographing requirement is single-ear photographing. Optionally, the photographing requirements can also be determined by the self-learning method of the user's behavior habits. For example, in the scene to be photographed in the gym, when the user manually switches the photographing requirement to single-ear photographing three consecutive times during photographing in the gym, the photographing requirement for this scene is determined to be single-ear photographing when this type of scene is recognized.

[0071] In this embodiment, after determining the scene to be photographed and the photographing requirements corresponding to the scene to be photographed, the photographing requirement for the corridor scene is single-ear photographing. At this time, the first earphone unit is controlled to execute a single-ear photographing action, while in the meeting room scene, the photographing requirement is double-ear photographing. At this time, the first earphone unit and the second earphone unit are controlled to execute the photographing action. Based on this, photographing control is performed through the photographing requirements corresponding to the scene to be photographed, improving the flexibility of photographing.

[0072] This embodiment provides a control method for earphones. When shooting based on the earphones, first, pre-shooting verification processing is performed through the first earphone unit. Subsequently, the collected environmental image is input into a pre-trained model for recognition, so as to determine the scene to be shot based on the recognition result. Finally, according to the shooting requirements associated with the to-be-shot scene, the earphones are controlled to perform a binaural shooting action or a monaural shooting action, avoiding the situation of high power consumption caused by using binaural shooting in a simple scene or poor shooting effect caused by using monaural shooting in a complex scene, and improving the flexibility of earphone shooting.

[0073] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , step S30 further includes steps S31 to S32:

[0074] Step S31, if the shooting requirement is binaural shooting, obtain the wearing information of the first earphone unit and the second earphone unit.

[0075] During binaural shooting, spatial audio recording such as 3D sound field and stereo ambient sound recording is usually required. At this time, the microphones of the left and right earphones need to work synchronously to capture the azimuth difference. If a single earphone is not worn, the microphone may be blocked, such as when the earphone does not enter the ear and touches the clothing, resulting in missing audio signals or reduced signal-to-noise ratio; at the same time, the binaural time difference and sound intensity difference require complete left and right channel data, otherwise the real sound source localization cannot be restored, such as when shooting in a meeting scene, it is impossible to distinguish between the left and right speakers based on the meeting recording; in addition, motion tracking is required during binaural shooting. If the earphone unit is not worn, it may be impossible to perform motion shooting tracking.

[0076] Therefore, in this embodiment, when the shooting requirement is binaural shooting, it is necessary to judge whether both earphones are in a worn state to improve the integrity of the shooting function and the user's shooting experience. It should be noted that the earphone can judge whether it is in a worn state through its own sensor, and the specific wearing state recognition process is not limited in this application.

[0077] Step S32, if both the first earphone unit and the second earphone unit are in a worn state, control the earphone to perform a binaural shooting action.

[0078] In this embodiment, when both earphone units of the earphone are in a worn state and binaural shooting is performed, two shooting instructions can be used to control the two earphone units to perform shooting actions simultaneously, or the shooting instruction can be sent from the first earphone unit to the second earphone unit.

[0079] Therefore, as an alternative implementation for controlling the earphones to perform a binaural shooting action, the image capture instruction can be synchronized to the second earphone unit based on the first earphone unit, and finally, the video modules of the first earphone unit and the second earphone unit are controlled to perform the image capture action. When the second earphone unit receives the image capture instruction, it will activate its video module. That is, after the second earphone unit receives the image capture instruction, it first checks the status of the camera device. If it is in an inactive state, it activates and turns on the camera device and then performs the shooting action. It can be understood that the video module of the first earphone unit is activated during the pre-shooting stage.

[0080] Optionally, when performing monaural shooting based on the earphones, the shooting action is directly performed through the video module of the first earphone unit to collect the current image information.

[0081] This embodiment provides a control method for earphones. When performing binaural shooting based on the earphones, it is determined whether the two earphone units are in a worn state, so as to ensure the shooting integrity and effectiveness when shooting based on the two earphone units and improve the shooting effect.

[0082] Based on the first embodiment or the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, after controlling the earphones to perform the binaural shooting action, the captured images can also be compared to determine whether there is a situation where the shooting of the two earphones is inconsistent. Therefore, after the step of controlling the earphones to perform the binaural shooting action, it is also necessary to obtain a first image and a second image. The first image and the second image are the images collected by the first earphone unit and the second earphone unit respectively. Subsequently, the overlapping area of the first graphic and the second image is determined, and feature matching is performed according to the feature information of the overlapping area to obtain the similarity of the overlapping area. If the similarity meets the preset similarity, the first image and the second image are synthesized to obtain the target image corresponding to the binaural shooting action. Otherwise, a prompt message for adjusting the wearing position of the earphone unit and / or inconsistent shooting is generated.

[0083] Specifically, during this process, the second earphone unit will send the picture to the first earphone unit, and the first earphone unit will compare the two pictures to check whether there are overlapping objects or people in the overlapping area of the pictures; if there is an overlap, it means that the shooting is successful, and then subsequent picture processing is performed; if there is no overlap, it is considered that the current user's wearing angle is incorrect, that is, there is a problem of inconsistent shooting time. At this time, a prompt message for inconsistent shooting can be output. In addition, the adjustment information of the wearing position can also be output to enable the user to adjust the wearing position information of the earphone.

[0084] Exemplarily, to help understand the implementation process of the control method of the earphones obtained by combining the above various embodiments, please refer toFigure 3 , Figure 3 A schematic flowchart of an alternative implementation of a method for controlling an earphone is provided. Specifically: After the user controls the first earphone to turn on the camera through gesture control or voice control, pre-shooting is performed based on the first earphone, and then the picture is compared with a pre-trained model to identify the quantity information of things, people, etc. in the picture. Then, a judgment is made according to the identified quantity and a pre-set threshold, so as to judge the complexity of the scene to be photographed based on the quantity information of the environmental image. If the complexity is high, that is, the quantity exceeds the preset threshold, enter the dual-ear photo-taking mode and execute the single-ear shooting action; if the complexity is low, that is, the quantity does not exceed the preset threshold, enter the single-ear photo-taking mode and execute the single-ear shooting action.

[0085] During the execution of the single-ear shooting action, first judge whether the current camera of the first earphone is in an active state. If it is in a sleep state, the first earphone executes the camera opening instruction and then takes a photo; if it is in an active state, directly execute the photo-taking instruction; it can be understood that during the shooting process, since pre-shooting needs to be performed through the first earphone, the first earphone is usually defaulted to be in an active state, that is, directly perform photo-taking processing based on the first earphone.

[0086] During the execution of the dual-ear shooting action, after the first earphone receives the camera opening instruction, it judges whether both earphones are worn and used, that is, whether they are both in a worn state. If so, the first earphone exchanges information through dual-ear connection, encapsulates the camera information in it, and then sends the camera opening instruction to the second earphone. If not, re-detection is performed until the user wears both earphones, or enter the single-ear photo-taking mode and execute the shooting action based on the first earphone. Then, after the second earphone receives the instruction from the first earphone, the second earphone checks the state of the current camera, whether it is in an active state. If not, activate and turn on the camera and then execute the photo-taking instruction; if so, directly execute the photo-taking instruction. At the same time, the first earphone directly executes the photo-taking instruction.

[0087] After both the first earphone and the second earphone have completed shooting, the second earphone sends the picture to the first earphone, and the first earphone compares the two pictures to check whether there are overlapping things or people in the overlapping area of the pictures; if there is an overlap, it means that the shooting is successful, and then subsequent picture processing is performed; if there is no overlap, it is considered that the current user's wearing angle is incorrect, resulting in a problem of inconsistent shooting times. At this time, the user is prompted by voice to indicate that the earphones are not worn properly and need to be re-worn before taking pictures again.

[0088] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be repeated in detail. On this basis, step S20, based on the image recognition result of the environment image, determines the step of the scene to be photographed and further includes steps S21 to S22:

[0089] Step S21, obtaining voice information collected by the earphone.

[0090] Step S22: determining the scene to be photographed according to the image recognition result and the voice information.

[0091] In this embodiment, in addition to determining the scene to be shot through the image recognition results, voice information in the current environment can also be collected through the microphone of the headset, and the voice information is analyzed and processed, so as to determine the scene to be shot by combining the voice information and the image recognition results, thereby improving the accuracy of identifying the scene to be shot, and further improving the accuracy of monaural shooting or binaural shooting based on the scene to be shot.

[0092] For example, in a conference shooting scene, when shooting with headphones, there are usually other participants communicating in the meeting, and the image recognition result is a display screen or a display curtain, so it can be determined that the scene to be shot is a conference scene. Or in a fitness scene, the image recognition result is the features of fitness equipment such as dumbbells and treadmill outlines, and there is a rapid breathing sound when running in the current environment, so it can be determined that the current scene to be shot is a fitness scene.

[0093] Optionally, in addition to determining the scene to be shot through voice information and image recognition results, the current motion data of the headset can also be combined for analysis and judgment, that is, first obtain the voice information and motion data collected by the headset, and then determine the scene to be shot based on the image recognition results, voice information and motion data.

[0094] For example, when a user is running on a track, the image recognition result is a single track information. At this time, it can be considered that the scene to be shot is a track with low complexity and the shooting is based on single-ear shooting. In the actual shooting process, the user is in a running state. In the state of severe running vibration, two earphones are usually required to shoot together to improve the shooting accuracy. Therefore, in order to improve the shooting accuracy and effectiveness, in addition to identifying the track information of the current environment, the user's breathing sound and the current changing position of the earphone are also collected. Based on this, it can be determined that the scene to be shot is a running scene.

[0095] This embodiment provides a control method for an earphone. When determining the scene to be photographed, auxiliary analysis and judgment are carried out by combining the voice information and motion data collected by the earphone, so as to improve the accuracy of the recognition of the scene to be photographed, avoid the earphone shooting based on the wrong shooting mode due to the error of the image recognition result, and improve the shooting effectiveness when shooting with the earphone based on the scene to be photographed.

[0096] This application provides an earphone, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the control method of the earphone in the above first embodiment.

[0097] The following refers to Figure 4 , which shows a schematic structural diagram of an earphone suitable for implementing the embodiments of the present application. Figure 4 The earphone shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0098] As Figure 4 shown, the earphone may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the earphone are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD, Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the earphone to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an earphone with various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems can be implemented or had.

[0099] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0100] The earphones provided by the present application adopt the earphone control method in the above-mentioned embodiment, and can solve the technical problem of poor flexibility in earphone shooting. Compared with the prior art, the beneficial effects of the earphones provided by the present application are the same as those of the earphone control method provided by the above-mentioned embodiment, and other technical features in the earphones are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0101] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0102] As mentioned above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0103] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the earphone control method in the above-mentioned embodiment.

[0104] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROMs), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0105] The above computer-readable storage medium may be included in the earphone; it may also exist separately and not be assembled into the earphone.

[0106] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the earphone, the earphone is caused to:

[0107] In response to a pre-shooting instruction, obtain the environmental image collected by the first earphone unit;

[0108] Determine the scene to be photographed according to the image recognition result of the environmental image;

[0109] Control the earphone to perform a single-ear shooting action or a two-ear shooting action according to the shooting requirements corresponding to the scene to be photographed.

[0110] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0112] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0113] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned control method of the earphone, and can solve the technical problem of poor flexibility in earphone shooting. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the control method of the earphone provided in the above embodiments, and will not be elaborated here.

[0114] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. A control method for an earphone, characterized in that, Applied to headphones, the headphones include a first headphone unit and a second headphone unit provided with a video module, and the control method of the headphones includes: Respond to a pre-shooting instruction, and obtain the environmental image collected by the first headphone unit; Determine the scene to be shot according to the image recognition result of the environmental image; Control the headphones to perform a single-ear shooting action or a double-ear shooting action according to the shooting requirements corresponding to the scene to be shot.

2. The control method of the earphone according to claim 1, characterized in that, The step of controlling the headphones to perform a single-ear shooting action or a double-ear shooting action according to the shooting requirements corresponding to the scene to be shot includes: If the shooting requirement is double-ear shooting, obtain the wearing information of the first headphone unit and the second headphone unit; If both the first headphone unit and the second headphone unit are in a worn state, control the headphones to perform a double-ear shooting action; Otherwise, control the first headphone unit to perform a single-ear shooting action.

3. The control method of the earphone according to claim 2, characterized in that, The step of controlling the headphones to perform a double-ear shooting action includes: Based on the first headphone unit, synchronize the image acquisition instruction to the second headphone unit. After receiving the image acquisition instruction, the second headphone unit activates the video module; Control the video modules of the first headphone unit and the second headphone unit to perform image acquisition actions.

4. The control method of the earphone according to any one of claims 2 or 3, characterized in that, After the step of controlling the headphones to perform a double-ear shooting action, it further includes: Obtain a first image and a second image, where the first image and the second image are the images collected by the first headphone unit and the second headphone unit respectively; Determine the overlapping area of the first image and the second image, and perform feature matching according to the feature information of the overlapping area to obtain a similarity; If the similarity meets the preset similarity, synthesize the first image and the second image to obtain the target image corresponding to the double-ear shooting action; Otherwise, generate a prompt message for the wearing position adjustment and / or inconsistent shooting of the headphone unit.

5. The control method of the earphone according to claim 1, characterized in that, Before the step of determining the scene to be shot according to the image recognition result of the environmental image, it further includes: Based on a pre-trained model, identify and process the object information of the environmental image to obtain the number of objects in the environmental image; Generate an image recognition result including the number of objects; The step of determining the scene to be shot according to the image recognition result of the environmental image includes: If the number of objects is greater than the preset number, determine that the scene to be shot is the first scene, and the shooting requirement of the first scene is double-ear shooting; Otherwise, determine that the scene to be shot is the second scene, and the shooting requirement of the second scene is single-ear shooting.

6. The control method of the earphone according to claim 1, characterized in that The step of determining the scene to be shot according to the image recognition result of the environmental image includes: Obtain the voice information collected by the headphones; Determine the scene to be shot according to the image recognition result and the voice information.

7. The control method of the earphone according to claim 6, wherein The step of determining the scene to be shot according to the environmental image recognition result further includes: Obtain the voice information and motion data collected by the headphones; Determine the scene to be shot according to the image recognition result, the voice information and the motion data.

8. The control method of the earphone according to claim 1, characterized in that, The shooting requirements include binaural shooting and monaural shooting, and the shooting requirements are determined based on any one of a preset instruction, the object quantity information of the image recognition result, or user behavior habit self-learning.

9. A headset, characterized in that, The earphone includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the control method of the earphone according to any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the control method of the earphone according to any one of claims 1 to 8 are implemented.