Earphone Wearing Detection Method and Device, Electronic Device, Storage Medium

Through the improvement of headphone wear detection methods, the use of audio data processing and logistic regression classification model, the problem of low accuracy of headphone wear detection in the prior art has been solved, and higher detection accuracy and stability have been achieved.

CN114333905BActive Publication Date: 2025-06-17SHENZHEN PHICOUSTIC SYST DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111519003.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2025-06-17
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

The detection accuracy of existing headphone wear detection methods is low, especially the detection effect of the distance measuring sensor is unstable, and the detection effect of the light sensor is limited in the night environment.

Method used

By obtaining the original audio data and the recorded audio data collected through the microphone, relevant calculation processing and logistic regression classification model operations are performed, scores are combined and judgment processing is performed to determine the wearing status of the headphones.

Benefits of technology

It improves the accuracy and stability of headphone wear detection, enhances the availability of detection in multiple scenarios, reduces the impact of noise on the detection process, and eliminates the process of parameter adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333905B_ABST
    Figure CN114333905B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method and device for detecting headphone wearing, an electronic device, and a storage medium, relating to the field of computer technology. The method for detecting headphone wearing includes: obtaining original audio data; obtaining recorded audio data collected by a microphone; performing correlation operation processing on the original audio data and the recorded audio data to obtain a correlation operation score; inputting the recorded audio data into a logistic regression classification model to obtain a model operation score; performing score merging processing on the correlation operation score and the model operation score to obtain a target score; and performing decision processing on the target score to obtain a headphone wearing detection result. The method for detecting headphone wearing provided by the embodiments of the present disclosure can improve the detection accuracy of headphone wearing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for detecting earphone wearing, an electronic device, and a storage medium. Background Art

[0002] The current method for detecting whether an earphone is worn is sensor detection, such as ranging sensor detection, light sensor detection, etc.; among them, the detection effect of the ranging sensor is unstable, and the detection effect of the light sensor has limitations in the night environment. Therefore, the detection accuracy of the sensor detection method is relatively low. Summary of the Invention

[0003] The main purpose of the embodiments of the present disclosure is to propose a method and device for detecting earphone wearing, an electronic device, and a storage medium, which can improve the detection accuracy of earphone wearing.

[0004] To achieve the above object, a first aspect of the embodiments of the present disclosure proposes a method for detecting earphone wearing, including:

[0005] Obtain original audio data;

[0006] Obtain recorded audio data collected by a microphone;

[0007] Perform correlation operation processing according to the original audio data and the recorded audio data to obtain a correlation operation score;

[0008] Input the recorded audio data into a logistic regression classification model to obtain a model operation score;

[0009] Perform score merging processing according to the correlation operation score and the model operation score to obtain a target score;

[0010] Perform decision processing on the target score to obtain an earphone wearing detection result.

[0011] In some embodiments, the performing correlation operation processing according to the original audio data and the recorded audio data to obtain a correlation operation score includes:

[0012] Obtain an autocorrelation score according to the original audio data;

[0013] Obtain a cross-correlation score according to the original audio data and the recorded audio data;

[0014] Normalize the cross-correlation score according to the autocorrelation score to obtain a correlation operation score.

[0015] In some embodiments, the obtaining an autocorrelation score according to the original audio data includes:

[0016] Obtained from a related preliminary table;

[0017] Perform a query process on the self - correlation preliminary table according to the original audio data to obtain the self - correlation score.

[0018] In some embodiments, obtaining the cross - correlation score according to the original audio data and the recorded audio data includes:

[0019] Perform a cross - correlation process on the original audio data and the recorded audio data to obtain a cross - correlation sequence; the cross - correlation sequence includes at least two cross - correlation parameters;

[0020] Use the cross - correlation parameter with the largest value in the cross - correlation sequence as the cross - correlation score.

[0021] In some embodiments, the method further includes training the logistic regression classification model, specifically including:

[0022] Obtain an initial network model; the initial network model includes a loss function gradient;

[0023] Obtain an audio data set; the audio data set includes first collected audio data in a state of wearing headphones and second collected audio data in a state of not wearing headphones;

[0024] Train the initial network model according to the audio data set and update the loss function gradient;

[0025] Update the initial network model according to the updated loss function gradient to obtain the logistic regression classification model.

[0026] In some embodiments, performing a score merging process on the relevant operation score and the model operation score to obtain a target score includes:

[0027] Obtain a weighting parameter;

[0028] Perform a weighted multiplication and addition process on the relevant operation score and the model operation score according to the weighting parameter to obtain the target score.

[0029] In some embodiments, performing a judgment process on the target score to obtain a headphone wearing detection result includes:

[0030] Obtain a judgment threshold;

[0031] Perform a judgment process on the target score according to the judgment threshold to obtain the headphone wearing detection result.

[0032] To achieve the above object, a second aspect of the present disclosure proposes a headphone wearing detection device, including:

[0033] An original audio data acquisition module, configured to acquire original audio data;

[0034] A recorded audio data acquisition module, configured to acquire recorded audio data collected by a microphone;

[0035] A related operation processing module, configured to perform related operation processing based on the original audio data and the recorded audio data to obtain a related operation score;

[0036] A classification model operation module, configured to input the recorded audio data into a logistic regression classification model to obtain a model operation score;

[0037] A score merging processing module, configured to perform score merging processing based on the related operation score and the model operation score to obtain a target score;

[0038] A decision module, configured to perform decision processing on the target score to obtain a headphone wearing detection result.

[0039] To achieve the above object, a third aspect of the present disclosure proposes an electronic device, including:

[0040] At least one memory;

[0041] At least one processor;

[0042] At least one program;

[0043] The program is stored in the memory, and the processor executes the at least one program to implement the method as described in the first aspect of the present disclosure.

[0044] To achieve the above object, a fourth aspect of the present disclosure proposes a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute:

[0045] The method as described in the first aspect above.

[0046] The headphone wearing detection method, device, electronic device, and storage medium proposed in the embodiments of the present disclosure first acquire original audio data, collect recorded audio data through a microphone, perform related operation processing based on the original audio data and the recorded audio data to obtain a related operation score, input the recorded audio data into a logistic regression classification model to obtain a model operation score, then perform score merging processing based on the related operation score and the model operation score to obtain a target score, and finally perform decision processing on the target score to obtain a headphone wearing detection result. Through the technical solution provided by the embodiments of the present disclosure, the stability of the headphone wearing detection process can be enhanced, the usability of the detection in multiple scenarios can be improved, and the detection accuracy of headphone wearing can be improved. Description of the Drawings

[0047] Figure 1 is a flowchart of the headphone wearing detection method provided by an embodiment of the present disclosure.

[0048] Figure 2 is Figure 1 a flowchart of step S130 in

[0049] Figure 3 is Figure 2 a flowchart of step S210 in

[0050] Figure 4 is Figure 2 a flowchart of step S220 in

[0051] Figure 5 is a partial flowchart of the headphone wearing detection method provided by another embodiment of the present disclosure.

[0052] Figure 6 is Figure 1 a flowchart of step S150 in

[0053] Figure 7 is Figure 1 a flowchart of step S160 in

[0054] Figure 8 is a block diagram of the headphone wearing detection device provided by an embodiment of the present disclosure.

[0055] Figure 9 is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present disclosure.

[0056] Reference numerals: original audio data acquisition module 810, recorded audio data acquisition module 820, correlation operation processing module 830, classification model operation module 840, score merging processing module 850, decision module 860, processor 901, memory 902, input / output interface 903, communication interface 904, bus 905. Detailed Embodiments

[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] It should be noted that although the functional modules are divided in the schematic diagram of the device and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division from that in the device or a different sequence from that in the flowchart. Terms such as "first" and "second" in the description, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used herein are only for the purpose of describing the embodiments of this invention and are not intended to limit this invention.

[0060] Detecting whether the earphone is worn is a very important function. Current methods for detecting whether the earphone is worn include sensor detection, such as ranging sensor detection, light sensor detection, etc.; among them, the detection effect of the ranging sensor is unstable, and the detection effect of the light sensor has limitations in the night environment. Therefore, the detection accuracy of the sensor detection method is relatively low.

[0061] In addition, current methods for detecting whether the earphone is worn also include the echo energy detection method, specifically implemented by playing a prompt tone and detecting the echo energy. Among them, the starting point is the time point when the prompt tone has energy. In the traditional echo energy detection method, it is very difficult to determine the starting point during the recording process, and it is necessary to repeatedly debug and capture the time point. In the echo energy detection method, the original audio for comparison calculates the data frame energy according to the starting point. If the starting point is detected inaccurately, there will be an error in the calculated value; the better the comparison points of the recorded sound and the comparison sound are aligned, the more accurate the calculation result will be.

[0062] The echo energy detection method requires parameter adjustment and has high requirements for the time starting point of the audio. Moreover, when the external noise is relatively strong, misjudgment may occur due to the equal energy between the sound source and the echo, resulting in relatively low detection efficiency and accuracy.

[0063] Based on this, embodiments of the present disclosure provide a method and apparatus for detecting headphone wearing, an electronic device, and a storage medium. First, by obtaining original audio data, the audio data is collected and recorded through a microphone. Based on the original audio data and the recorded audio data, relevant arithmetic processing is performed to obtain a relevant arithmetic score. The recorded audio data is input into a logistic regression classification model to obtain a model arithmetic score. Then, based on the relevant arithmetic score and the model arithmetic score, score merging processing is performed to obtain a target score. Finally, judgment processing is performed on the target score to obtain a headphone wearing detection result. Through the technical solution provided by the embodiments of the present disclosure, the stability of the headphone wearing detection process can be enhanced, the usability of the detection in multiple scenarios can be improved, the process of parameter tuning is omitted, the influence of noise on the detection process is reduced, the dependence on the accuracy of the audio time start point is reduced, and the detection efficiency and accuracy of headphone wearing are improved.

[0064] Embodiments of the present disclosure provide a method and apparatus for detecting headphone wearing, an electronic device, and a storage medium, which will be specifically described through the following embodiments. First, the method for detecting headphone wearing in the embodiments of the present disclosure will be described.

[0065] The method for detecting headphone wearing provided by the embodiments of the present disclosure relates to the field of computer technology, and particularly relates to the field of audio collection and processing technology. The method for detecting headphone wearing provided by the embodiments of the present disclosure can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc.; the server can be an independent server, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; the software can be an application implementing the method for detecting headphone wearing, etc., but is not limited to the above forms.

[0066] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0067] An embodiment of the present disclosure provides a method for detecting headphone wearing, including: obtaining original audio data; obtaining recorded audio data collected by a microphone; performing correlation operation processing based on the original audio data and the recorded audio data to obtain a correlation operation score; inputting the recorded audio data into a logistic regression classification model to obtain a model operation score; performing score merging processing based on the correlation operation score and the model operation score to obtain a target score; and performing a decision process on the target score to obtain a headphone wearing detection result.

[0068] Figure 1 is an optional flowchart of the method for detecting headphone wearing provided by the embodiment of the present disclosure. Figure 1 The method in may include but is not limited to steps S110 to S160, specifically including:

[0069] S110, obtaining original audio data;

[0070] S120, obtaining recorded audio data collected by a microphone;

[0071] S130, performing correlation operation processing based on the original audio data and the recorded audio data to obtain a correlation operation score;

[0072] S140, inputting the recorded audio data into a logistic regression classification model to obtain a model operation score;

[0073] S150, performing score merging processing based on the correlation operation score and the model operation score to obtain a target score;

[0074] S160, performing a decision process on the target score to obtain a headphone wearing detection result.

[0075] It should be noted that Figure 1The headphone wearing detection method shown is used for headphones with a built-in microphone, such as over-ear headphones, etc. Among them, the built-in microphone is a functional microphone, and its function is to collect audio to complete the detection of the headphone wearing situation.

[0076] In step S110, the original audio data is a preset audio for wearing detection, also known as a detection prompt tone. The volume of the detection prompt tone is usually 50 dB or above. The present disclosure does not limit the size and format of the detection prompt tone, which can be user-defined as long as it can achieve audio collection and judgment.

[0077] In step S120, the recorded audio data is the audio collected by the headphone through the microphone. This audio is obtained by the headphone's speaker playing the original audio data and then being collected by the microphone, and is used for subsequent operations and judgment processes.

[0078] In some embodiments, the number of microphones is at least one. When there are multiple microphones, the recorded audio data obtained by multiple microphones can be superimposed and then subjected to mean smoothing processing to improve the quality of the recorded audio data, thereby improving the accuracy of subsequent judgment.

[0079] In step S130, the correlation operation processing includes autocorrelation processing and cross-correlation processing. Among them, the object of autocorrelation processing is the original audio data, and the object of cross-correlation processing is the original audio data and the recorded audio data. By using the properties of audio autocorrelation and cross-correlation, the correlation between the original audio data and the recorded audio data is evaluated to obtain a correlation operation score.

[0080] Specifically, the correlation operation score represents the similarity degree of the audio before and after collection. The larger the correlation operation score, the more similar the audio before and after collection, indicating a greater possibility that the headphones are in the worn state; conversely, the smaller the correlation operation score, the less similar the audio before and after collection, indicating a greater possibility that the headphones are in the unworn state.

[0081] In step S140, the logistic regression classification model is a binary classification neural network model, which is obtained by training an initial network model with a dataset and is used to classify the recorded audio data. Specifically, after the recorded audio data is input into the logistic regression classification model, the model outputs a model operation score. The model operation score is a normalized score, and its value range is (0, 1).

[0082] Specifically, the model operation score represents the possibility of being in the worn state. The larger the model operation score, the greater the possibility that the headphones are in the worn state; conversely, the smaller the model operation score, the smaller the possibility that the headphones are in the worn state.

[0083] In step S150, the score merging process is specifically as follows: The obtained relevant operation score and the model operation score are merged to obtain the target score. The score merging process enables the target score to have the comprehensive characteristics of the two evaluation methods, thereby making the accuracy of headphone wearing detection higher.

[0084] Specifically, the methods of score merging process include, but are not limited to: multiplication, addition, weighted multiplication and addition, etc.

[0085] In step S160, the decision-making process is specifically as follows: The obtained target score is compared with the decision threshold to obtain the headphone wearing detection result. Among them, the headphone wearing detection result includes: the headphone is worn, the headphone is not worn.

[0086] The headphone wearing detection method provided by the embodiments of the present disclosure first obtains the original audio data by collecting and recording the audio data through a microphone, performs relevant operation processing according to the original audio data and the recorded audio data to obtain the relevant operation score, inputs the recorded audio data into the logistic regression classification model to obtain the model operation score, then performs score merging processing according to the relevant operation score and the model operation score to obtain the target score, and finally performs decision-making processing on the target score to obtain the headphone wearing detection result. Through the technical solution provided by the embodiments of the present disclosure, the stability of the headphone wearing detection process can be enhanced, the usability of the detection in multiple scenarios can be improved, the process of parameter adjustment can be omitted, the influence of noise on the detection process can be reduced, the dependence on the accuracy of the audio time starting point can be reduced, and the detection efficiency and accuracy of headphone wearing can be improved.

[0087] In some embodiments, performing relevant operation processing according to the original audio data and the recorded audio data to obtain the relevant operation score includes: obtaining the autocorrelation score according to the original audio data; obtaining the cross-correlation score according to the original audio data and the recorded audio data; performing normalization processing on the cross-correlation score according to the autocorrelation score to obtain the relevant operation score.

[0088] Figure 2 It is a flowchart of step S130 in some embodiments. Figure 2 The schematic step S130 includes, but is not limited to, steps S210 to S230:

[0089] S210, obtaining the autocorrelation score according to the original audio data;

[0090] S220, obtaining the cross-correlation score according to the original audio data and the recorded audio data;

[0091] S230, performing normalization processing on the cross-correlation score according to the autocorrelation score to obtain the relevant operation score.

[0092] In step 210, the original audio data is a preset audio for wearing detection, also known as a detection prompt tone; the autocorrelation score characterizes the autocorrelation of the original audio data and is used for normalizing the cross-correlation score.

[0093] In step 220, the recorded audio data is the audio collected by the earphone through the microphone, which is obtained by the microphone after the original audio data is played by the earphone's speaker; the cross-correlation score characterizes the cross-correlation between the original audio data and the recorded audio data and is used to generate the cross-correlation score.

[0094] In step 230, the normalization process is specifically as follows: using the autocorrelation score as the denominator and the cross-correlation score as the numerator for calculation to obtain a correlation operation score. The correlation operation score obtained by the normalization process can better reflect the correlation degree of the voice before and after collection, which is more beneficial for the detection result.

[0095] Specifically, the normalization process is: corr = corr1 / corr2, where corr1 is the cross-correlation score, corr2 is the autocorrelation score, and corr is the correlation operation score obtained by the normalization process.

[0096] In some embodiments, obtaining the autocorrelation score according to the original audio data includes: obtaining an autocorrelation preparation table; performing a query process in the autocorrelation preparation table according to the original audio data to obtain the autocorrelation score.

[0097] Figure 3 It is a flowchart of step S210 in some embodiments. Figure 3 The schematic step S210 includes but is not limited to steps S310 to S320:

[0098] S310, obtaining an autocorrelation preparation table;

[0099] S320, performing a query process in the autocorrelation preparation table according to the original audio data to obtain the autocorrelation score.

[0100] In step S310, the autocorrelation preparation table is a record of the autocorrelation scores of the original audio data prepared in advance. Specifically, first perform an autocorrelation process on the original audio data to obtain an autocorrelation sequence; the autocorrelation sequence includes at least two cross-correlation parameters, and then use the largest autocorrelation parameter in the autocorrelation sequence as the autocorrelation score, denoted as corr2, for subsequent normalization.

[0101] In step S320, since the original audio data does not need to be collected in real time, an autocorrelation preparation table is set up to obtain the autocorrelation score by means of a query process, so as to improve the detection efficiency of the earphone wearing detection method.

[0102] In some embodiments, obtaining a cross-correlation score based on the original audio data and the recorded audio data includes: performing cross-correlation processing on the original audio data and the recorded audio data to obtain a cross-correlation sequence; the cross-correlation sequence includes at least two cross-correlation parameters; using the cross-correlation parameter with the largest value in the cross-correlation sequence as the cross-correlation score.

[0103] Figure 4 is a flowchart of step S220 in some embodiments, Figure 4 The illustrated step S220 includes but is not limited to steps S410 to S420:

[0104] S410, performing cross-correlation processing on the original audio data and the recorded audio data to obtain a cross-correlation sequence;

[0105] S420, using the cross-correlation parameter with the largest value in the cross-correlation sequence as the cross-correlation score.

[0106] In step S410, the cross-correlation sequence includes at least two cross-correlation parameters, where the cross-correlation parameter is an element of the cross-correlation sequence.

[0107] Specifically, if the original audio data is A = [1 2 3] and the recorded audio data is B = [2, 3, 1], then the cross-correlation processing process is: C = conv(A, B) = [2 7 13 11 3], where C = [2 7 13 11, 3] is the obtained cross-correlation sequence.

[0108] In step S420, if the cross-correlation sequence is C = [2 7 13 11 3], including 5 cross-correlation parameters, which are 2, 7, 13, 11, and 3 respectively, and the largest value is 13, then 13 is used as the cross-correlation score, denoted as corr1.

[0109] In some embodiments, the method further includes training a logistic regression classification model, specifically including: obtaining an initial network model; the initial network model includes a loss function gradient; obtaining an audio data set; the audio data set includes first collected audio data in the state of wearing headphones and second collected audio data in the state of not wearing headphones; training the initial network model according to the audio data set to update the loss function gradient; updating the initial network model according to the updated loss function gradient to obtain a logistic regression classification model.

[0110] As Figure 5 shown, Figure 5 is a flowchart of a headphone wearing detection method provided by some other embodiments. The headphone wearing detection method further includes:

[0111] S510, obtaining an initial network model;

[0112] S520, obtaining an audio data set;

[0113] S530. Train the initial network model according to the audio data set and update the gradient of the loss function.

[0114] S540. Update the initial network model according to the updated gradient of the loss function to obtain a logistic regression classification model.

[0115] In step S510, the initial network model includes the gradient of the loss function, which is used to train the initial network model to determine the values of the model parameters so as to obtain a logistic regression classification model.

[0116] In step S520, the audio data set includes the first collected audio data in the state of wearing headphones and the second collected audio data in the state of not wearing headphones.

[0117] In steps S530 to S540, through the training of the audio data set, a logistic regression classification model is obtained. The logistic regression classification model is a binary classification neural network model, which is obtained by training the initial network model with the data set and is used to classify the recorded audio data.

[0118] In some embodiments, score merging processing is performed on the relevant operation score and the model operation score to obtain a target score, including: obtaining a weighting parameter; performing weighted multiplication and addition processing on the relevant operation score and the model operation score according to the weighting parameter to obtain the target score.

[0119] Figure 6 It is a flowchart of step S150 in some embodiments. Figure 6 The schematic step S150 includes but is not limited to steps S610 to S620:

[0120] S610. Obtain a weighting parameter.

[0121] S620. Perform weighted multiplication and addition processing on the relevant operation score and the model operation score according to the weighting parameter to obtain the target score.

[0122] It should be noted that the methods of score merging processing include but are not limited to: multiplication, addition, weighted multiplication and addition, etc. Figure 6 The embodiment shown adopts the method of weighted multiplication and addition.

[0123] In step S610, the weighting parameter is the parameter used for weighted multiplication and addition.

[0124] In step S620, if the weighting parameter is λ, the weighted multiplication and addition processing process is: y = λ * x + (1 - λ) * corr; where y is the obtained target score, x is the model operation score, and corr is the relevant operation score.

[0125] It should be noted that the weighting parameter λ represents the degree of bias towards a certain calculation method. Therefore, the weighting parameter λ can be adjusted according to the actual situation.

[0126] In some embodiments, a determination process is performed on the target score to obtain an earphone wearing detection result, including: obtaining a determination threshold; performing a determination process on the target score according to the determination threshold to obtain an earphone wearing detection result.

[0127] Figure 7 It is a flowchart of step S160 in some embodiments. Figure 7 The schematic step S160 includes but is not limited to steps S710 to S720:

[0128] S710, obtain a determination threshold;

[0129] S720, perform a determination process on the target score according to the determination threshold to obtain an earphone wearing detection result.

[0130] In step S710, if the method of score merging processing is multiplication, that is, y = x * corr, then the empirical value of the determination threshold is 0.7.

[0131] In step S720, if the calculated target score y is greater than or equal to 0.7, the earphone wearing detection result is that the earphone is worn; if the calculated target score y is less than 0.7, the earphone wearing detection result is that the earphone is not worn. It should be noted that in the case of earphone wearing, the envelope similarity between the recorded audio data recorded by the microphone and the original audio data is extremely high. Therefore, its correlation is stronger, and the target score y shown is larger.

[0132] Please refer to Figure 8 , Figure 8 illustrates an earphone wearing detection device according to an embodiment. The earphone wearing detection device includes: an original audio data acquisition module 810, a recorded audio data acquisition module 820, a correlation operation processing module 830, a classification model operation module 840, a score merging processing module 850, and a determination module 860.

[0133] Among them, the original audio data acquisition module 810 is used to acquire original audio data; the recorded audio data acquisition module 820 is used to acquire recorded audio data collected by a microphone; the relevant operation processing module 830 is used to perform relevant operation processing on the original audio data and the recorded audio data to obtain a relevant operation score; the classification model operation module 840 is used to input the recorded audio data into a logistic regression classification model to obtain a model operation score; the score merging processing module 850 is used to perform score merging processing on the relevant operation score and the model operation score to obtain a target score; the decision module 860 is used to perform decision processing on the target score to obtain a headphone wearing detection result.

[0134] The specific implementation manner of the headphone wearing detection device in this embodiment is basically the same as that of the above headphone wearing detection method, belonging to the same inventive concept, and will not be elaborated here.

[0135] This disclosure embodiment also provides an electronic device, including:

[0136] At least one memory;

[0137] At least one processor;

[0138] At least one program;

[0139] The program is stored in the memory, and the processor executes the at least one program to implement the headphone wearing detection method described above in this disclosure. This electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, abbreviated as PDA), an in-vehicle computer, etc.

[0140] Please refer to Figure 9 , Figure 9 which illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0141] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this disclosure embodiment;

[0142] The memory 902 can be implemented in the form of a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and the processor 901 is used to call and execute the headphone wearing detection method of the embodiments of the present disclosure;

[0143] The input / output interface 903 is used to implement information input and output;

[0144] The communication interface 904 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and

[0145] The bus 905 transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0146] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.

[0147] The embodiments of the present disclosure also provide a storage medium. This storage medium is a computer-readable storage medium, and this computer-readable storage medium stores computer-executable instructions. These computer-executable instructions are used to cause a computer to execute the above-mentioned headphone wearing detection method.

[0148] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0149] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0150] Those skilled in the art can understand that Figure 1-7 the technical solutions shown in

[0151] do not constitute a limitation to the embodiments of the present disclosure, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0152] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0153] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0154] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0155] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0156] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0158] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0159] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present disclosure. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present disclosure shall fall within the scope of the rights of the embodiments of the present disclosure.

Claims

1. A method for detecting headphone wearing, characterized in that, Including: Obtain original audio data; Obtain recorded audio data collected by a microphone; Perform relevant arithmetic processing based on the original audio data and the recorded audio data to obtain a relevant arithmetic score; Obtain an initial network model; the initial network model includes a loss function gradient; Obtain an audio data set; the audio data set includes first collected audio data in a state of wearing headphones and second collected audio data in a state of not wearing headphones; Train the initial network model according to the audio data set and update the loss function gradient; Update the initial network model according to the updated loss function gradient to obtain a logistic regression classification model; Input the recorded audio data into the logistic regression classification model to obtain a model arithmetic score; Perform score merging processing based on the relevant arithmetic score and the model arithmetic score to obtain a target score; Perform a judgment process on the target score to obtain a headphone wearing detection result.

2. The method according to claim 1, characterized in that, The performing relevant arithmetic processing based on the original audio data and the recorded audio data to obtain a relevant arithmetic score includes: Obtain an autocorrelation score according to the original audio data; Obtain a cross-correlation score according to the original audio data and the recorded audio data; Normalize the cross-correlation score according to the autocorrelation score to obtain a relevant arithmetic score.

3. The method according to claim 2, characterized in that, The obtaining an autocorrelation score according to the original audio data includes: Obtain an autocorrelation preparatory table; Perform a query process on the autocorrelation preparatory table according to the original audio data to obtain the autocorrelation score.

4. The method according to claim 2, characterized in that, The obtaining a cross-correlation score according to the original audio data and the recorded audio data includes: Perform a cross-correlation process on the original audio data and the recorded audio data to obtain a cross-correlation sequence; the cross-correlation sequence includes at least two cross-correlation parameters; Use the cross-correlation parameter with the largest value in the cross-correlation sequence as the cross-correlation score.

5. The method according to claim 1, characterized in that, The performing score merging processing based on the relevant arithmetic score and the model arithmetic score to obtain a target score includes: Obtain a weighting parameter; Perform a weighted multiplication and addition process on the relevant arithmetic score and the model arithmetic score according to the weighting parameter to obtain the target score.

6. The method according to any one of claims 1 to 5, characterized in that, The performing a judgment process on the target score to obtain a headphone wearing detection result includes: Obtain a judgment threshold; Perform a judgment process on the target score according to the judgment threshold to obtain the headphone wearing detection result.

7. A headphone wearing detection device, characterized in that, Including: An original audio data acquisition module for obtaining original audio data; A recorded audio data acquisition module for obtaining recorded audio data collected by a microphone; A relevant arithmetic processing module for performing relevant arithmetic processing based on the original audio data and the recorded audio data to obtain a relevant arithmetic score; A model training module for obtaining an initial network model; the initial network model includes a loss function gradient; obtaining an audio data set; the audio data set includes first collected audio data in a state of wearing headphones and second collected audio data in a state of not wearing headphones; Training the initial network model according to the audio data set, and updating the gradient of the loss function; updating the initial network model according to the updated gradient of the loss function to obtain a logistic regression classification model; A classification model operation module, configured to input the recorded audio data into the logistic regression classification model to obtain a model operation score; A score merging processing module, configured to perform score merging processing according to the relevant operation score and the model operation score to obtain a target score; A decision module, configured to perform decision processing on the target score to obtain a headphone wearing detection result.

8. An electronic device, characterized in that, Including: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes the at least one program to implement: The method according to any one of claims 1 to 6.

9. A storage medium, the storage medium being a computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute: The method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Volume adjusting method and mobile terminal

    CN106027809A

  • Wearing situation detection method and device of wireless headset, and wireless headset

    CN108810788A