Mask wearing detection methods, devices, storage media and electronic equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]为克服相关技术中存在的问题,本公开提供一种口罩佩戴检测方法、装置、存储介质及电子设备,以解决口罩佩戴检测成本较高,且无法针对所有人进行检测的问题
[0060] When the communication device is in a call, the correlation coefficient between the sound signal from the transmitting device and the sound signal received by the receiving device is determined, along with the distance between the transmitting object and the communication device. Since the sound signal received by the receiving device includes the reflected sound signal from the reflecting object, and the distance between the reflecting object and the communication device, as well as whether the reflecting object is wearing a mask, will affect the correlation coefficient, it is possible to determine whether the reflecting object is wearing a mask based on the calculated correlation coefficient and the determined distance. In this way, the mask-wearing status of users can be detected using the existing sound signal of the communication device. This allows for one-to-one mask-wearing detection for each user of each communication device, preventing missed detections. Furthermore, it utilizes existing sound resources without requiring additional devices or signals, reducing resource waste.
Smart Images

Figure CN116840843B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, and in particular to a method, apparatus, storage medium, and electronic device for detecting mask wearing. Background Technology
[0002] Respiratory infectious diseases are mainly transmitted through droplets, but can also be transmitted through direct or indirect contact. Pathogens invade the body through the respiratory tract, including the nasal cavity, pharynx, trachea, and bronchi. Wearing masks can effectively reduce the spread of respiratory infectious viruses. Meanwhile, wearing masks in public places has become a necessary means of routine epidemic prevention and control. However, manually assessing mask-wearing requires a significant amount of time and manpower, resulting in low efficiency.
[0003] In related technologies, mask-wearing detection typically involves training a neural network model to detect and recognize facial images scanned by physical instruments. This method requires a large number of training samples for model training, making the training process complex, and the physical instruments are generally expensive. Furthermore, in crowded places, it may be impossible to scan the facial images of everyone, making it impossible to detect mask-wearing on everyone. Summary of the Invention
[0004] To overcome the problems existing in related technologies, this disclosure provides a method, device, storage medium and electronic device for mask wearing detection, so as to solve the problems that mask wearing detection is costly and cannot be tested for everyone.
[0005] According to a first aspect of the present disclosure, a method for detecting mask wearing is provided, comprising:
[0006] When the communication device is in a call state, a first sound signal emitted by the sound-emitting device of the communication device and a second sound signal received by the sound-receiving device of the communication device are determined. The second sound signal includes the reflected sound signal of the first sound signal after being reflected by a reflective object and reaching the sound-receiving device.
[0007] Determine the correlation coefficient between the second sound signal and the first sound signal;
[0008] Determine the distance between the reflective object and the communication device;
[0009] If the calling device has a pre-stored target correspondence that matches the distance and the correlation coefficient, it is determined that the user of the calling device is wearing a mask. The target correspondence is the correspondence between the correlation coefficient between the first sound signal and the second sound signal measured when the user is wearing a mask and holding the calling device to make a call, and the distance between the user's face and the calling device.
[0010] Optionally, the second sound signal further includes the direct sound signal from the first sound signal reaching the receiving device, and the method further includes:
[0011] Record the sound signal emitted by the sound-producing device and the sound signal received by the sound-receiving device to obtain the first sound signal and the second sound signal;
[0012] The first sound signal recorded within a preset time period is processed into frames to obtain a first sound frame sequence, and the second sound signal is processed into frames to obtain a second sound frame sequence.
[0013] Based on the first sound frame sequence and the second sound frame sequence, calculate the correlation coefficient for each sound frame in the second sound frame sequence in sequence.
[0014] The determination includes the correlation coefficient between the second sound signal and the first sound signal, including:
[0015] The second peak of the calculated correlation coefficient is taken as the correlation coefficient between the second sound signal and the first sound signal.
[0016] Optionally, determining the distance between the reflective object and the communication device includes:
[0017] Determine the time difference between the recording times of the audio frames corresponding to the first peak and the second peak of the calculated correlation coefficient;
[0018] The distance between the reflecting object and the communication device is determined based on the time difference and the speed of sound propagation.
[0019] Optionally, the step of sequentially calculating the correlation coefficient for each sound frame in the second sound frame sequence based on the first sound frame sequence and the second sound frame sequence includes:
[0020] For each sound frame in the second sound frame sequence, a target sound frame sequence corresponding to that sound frame is determined, wherein the number of frames in the target sound frame sequence is the same as that in the first sound frame sequence;
[0021] The correlation coefficient between the target sound frame sequence and the first sound frame sequence is used as the correlation coefficient for the corresponding sound frame.
[0022] Optionally, before determining the first sound signal emitted by the speaking device of the communication device and the second sound signal received by the receiving device of the communication device, the method further includes:
[0023] It is determined that the calling device is in a handheld calling state.
[0024] Optionally, it also includes:
[0025] If the calling device does not have a pre-stored target correspondence that matches the distance and the correlation coefficient, it is determined that the user of the calling device is not wearing a mask.
[0026] Optionally, it also includes:
[0027] If it is determined that the user of the calling device is not wearing a mask, the user of the calling device will be prompted to wear a mask.
[0028] Optionally, it also includes:
[0029] If it is determined that the user of the calling device is wearing a mask, the sound signal received by the user from the sound receiver of the calling device is compensated, and the compensated sound signal is sent to the other party in the call.
[0030] According to a second aspect of the present disclosure, a mask-wearing detection device is provided, comprising:
[0031] The first determining module is configured to, when the calling device is in a calling state, determine a first sound signal emitted by the speaking device of the calling device and a second sound signal received by the receiving device of the calling device, wherein the second sound signal includes the reflected sound signal of the first sound signal after being reflected by a reflecting object and reaching the receiving device.
[0032] The second determining module is configured to determine the correlation coefficient between the second sound signal and the first sound signal;
[0033] The third determining module is configured to determine the distance between the reflecting object and the communication device;
[0034] The fourth determining module is configured to determine that the user of the calling device is wearing a mask when the calling device has a target correspondence that matches the distance and the correlation coefficient. The target correspondence is the correspondence between the correlation coefficient between the first sound signal and the second sound signal measured when the user is wearing a mask and holding the calling device to make a call, and the distance between the user's face and the calling device.
[0035] Optionally, the second sound signal further includes the direct sound signal from the first sound signal reaching the receiving device, and the device further includes:
[0036] The recording module is configured to record the sound signal emitted by the sound-producing device and the sound signal received by the sound-receiving device, thereby obtaining the first sound signal and the second sound signal;
[0037] The frame segmentation module is configured to perform frame segmentation processing on the first sound signal recorded within a preset time period to obtain a first sound frame sequence, and to perform frame segmentation processing on the second sound signal to obtain a second sound frame sequence.
[0038] The calculation module is configured to calculate the correlation coefficient of each sound frame in the second sound frame sequence in turn based on the first sound frame sequence and the second sound frame sequence.
[0039] The second determining module is configured as follows:
[0040] The second peak of the calculated correlation coefficient is taken as the correlation coefficient between the second sound signal and the first sound signal.
[0041] Optionally, the third determining module is configured as follows:
[0042] Determine the time difference between the recording times of the audio frames corresponding to the first peak and the second peak of the calculated correlation coefficient;
[0043] The distance between the reflecting object and the communication device is determined based on the time difference and the speed of sound propagation.
[0044] Optionally, the computing module is configured as follows:
[0045] For each sound frame in the second sound frame sequence, a target sound frame sequence corresponding to that sound frame is determined, wherein the number of frames in the target sound frame sequence is the same as that in the first sound frame sequence;
[0046] The correlation coefficient between the target sound frame sequence and the first sound frame sequence is used as the correlation coefficient for the corresponding sound frame.
[0047] Optionally, the device further includes a fifth determining module, which is configured to:
[0048] Before determining the first sound signal emitted by the voice output device of the calling device and the second sound signal received by the voice receiver of the calling device, it is determined that the calling device is in a handheld calling state.
[0049] Optionally, the device further includes a sixth determining module, the sixth determining module being configured to:
[0050] If the calling device does not have a pre-stored target correspondence that matches the distance and the correlation coefficient, it is determined that the user of the calling device is not wearing a mask.
[0051] Optionally, the device further includes a prompting module, which is configured to:
[0052] If it is determined that the user of the calling device is not wearing a mask, the user of the calling device will be prompted to wear a mask.
[0053] Optionally, the device further includes a compensation module, which is configured to:
[0054] If it is determined that the user of the calling device is wearing a mask, the sound signal received by the user from the sound receiver of the calling device is compensated, and the compensated sound signal is sent to the other party in the call.
[0055] According to a third aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the mask-wearing detection method provided in the first aspect of the present disclosure.
[0056] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising:
[0057] A memory on which computer programs are stored;
[0058] The processor is configured to execute the computer program in the memory to implement the steps of the mask-wearing detection method provided in the first aspect of this disclosure.
[0059] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0060] When the communication device is in a call, the correlation coefficient between the sound signal from the transmitting device and the sound signal received by the receiving device is determined, along with the distance between the transmitting object and the communication device. Since the sound signal received by the receiving device includes the reflected sound signal from the reflecting object, and the distance between the reflecting object and the communication device, as well as whether the reflecting object is wearing a mask, will affect the correlation coefficient, it is possible to determine whether the reflecting object is wearing a mask based on the calculated correlation coefficient and the determined distance. In this way, the mask-wearing status of users can be detected using the existing sound signal of the communication device. This allows for one-to-one mask-wearing detection for each user of each communication device, preventing missed detections. Furthermore, it utilizes existing sound resources without requiring additional devices or signals, reducing resource waste.
[0061] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0062] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0063] Figure 1 This is a flowchart illustrating a mask-wearing detection method according to an exemplary embodiment.
[0064] Figure 2 This is a schematic diagram illustrating the amplitude of a sound signal according to another exemplary embodiment.
[0065] Figure 3 This is a schematic diagram illustrating a correlation coefficient peak according to an exemplary embodiment.
[0066] Figure 4 This is a flowchart illustrating a mask-wearing detection method according to another exemplary embodiment.
[0067] Figure 5 This is a block diagram illustrating a mask-wearing detection device according to an exemplary embodiment.
[0068] Figure 6 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0069] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0070] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.
[0071] Figure 1 This is a flowchart illustrating a mask-wearing detection method according to an exemplary embodiment, such as... Figure 1 As shown, the mask-wearing detection method used in the terminal includes the following steps:
[0072] In step S11, when the communication device is in a call state, a first sound signal emitted by the voice device of the communication device and a second sound signal received by the voice receiver of the communication device are determined. The second sound signal includes the reflected sound signal of the first sound signal after being reflected by a reflective object and reaching the voice receiver.
[0073] It should be understood that a call can be a call made using a circuit-switched communication network, i.e., a call made by dialing a phone number, or a call made using a packet-switched internet network, such as a call made using social software such as WeChat or QQ, or a call where the voice and receiver of other call devices are working simultaneously. This disclosure does not limit the scope of the call.
[0074] When the communication device is in a call state, both the sound output device and the receiver are operational. The sound signal received by the receiver includes the sound signal emitted by the sound output device, which is reflected by a reflective object, such as the user's face, before reaching the receiver.
[0075] It should also be understood that, during a call, the calling device periodically samples and quantizes the sound emitted by its speaker and received by its receiver to obtain a result such as... Figure 2 The audio signal amplitude corresponding to each sampling point is shown. This disclosure does not limit the sampling and quantization methods.
[0076] In step S12, the correlation coefficient between the second sound signal and the first sound signal is determined.
[0077] It should be understood that the correlation coefficient is used to characterize the similarity between the second audio signal and the first audio signal. Considering that both the first and second audio signals can include multiple audio frames, in one possible implementation, the correlation coefficient between the first and second audio signals can be the average of the correlation coefficients between the individual frames in the audio signal. In another possible implementation, considering that the correlation coefficient of the audio frames will peak when the second audio signal received by the receiver of the calling device includes reflected audio signals, the peak value of the correlation coefficient between the audio frames after the reflected audio signal reaches the receiver can also be determined and used to characterize the correlation coefficient between the first and second audio signals. Figure 3As shown, X represents the first audio signal, Y represents the second audio signal, and the peak value of the correlation coefficient is delayed by 151 sampling points. This means that the second audio signal is received by the receiver of the communication device 151 points after the first audio signal is emitted by the device's output device. The received audio signal is the reflected sound signal from the first audio signal after reflection by a reflective object. The specific delay time needs to be determined based on the audio signal's sampling rate (FS). If the sampling rate is 8kHz, the delay of the second audio signal is 151 / 8000 = 0.0189 seconds.
[0078] In step S13, the distance between the reflective object and the communication device is determined.
[0079] In step S14, if the calling device has a pre-stored target correspondence relationship that matches the distance and correlation coefficient, it is determined that the user of the calling device is wearing a mask. The target correspondence relationship is the correspondence relationship between the correlation coefficient between the first sound signal and the second sound signal measured when the user is wearing a mask and holding the calling device to make a call, and the distance between the user's face and the calling device.
[0080] It should be understood that, due to the different absorption rates of the reflecting objects (wearing masks and not wearing masks), the absorption rate of the first sound signal varies. Therefore, when the reflecting object and the communication device are at the same distance, whether the reflecting object is wearing a mask will affect the calculated correlation coefficient between the first and second sound signals. Thus, for different distances, the correlation coefficient corresponding to the distance when the user is wearing a mask can be pre-calibrated, thereby obtaining the correspondence between distance and correlation coefficient. Furthermore, since different mask materials (e.g., non-woven fabric, cotton, sponge, etc.) also have different absorption rates of the first sound signal, the correlation coefficient between the first and second sound signals will also differ when the reflecting object and the communication device are at the same distance and the reflecting object is wearing masks of different materials. Therefore, one possible implementation of this disclosure is to pre-calibrate the correspondence between distance and correlation coefficient when the user is wearing masks of different materials, thereby obtaining multiple sets of one-to-one correspondences with mask materials, where each set of correspondences may also include correspondences for the same mask material at different distances.
[0081] Therefore, after measuring the distance between the reflecting object and the communication device and the correlation coefficient of the first and second sound signals, step S14 can determine whether there is a correspondence between the measured distance and correlation coefficient in the pre-stored correspondence between distance and correlation coefficient. If there is, it can be determined that the user of the communication device is wearing a mask.
[0082] For the measured correspondence (i.e., the correspondence between the measured current distance and the calculated correlation coefficient), it is determined whether there exists a target correspondence that matches the measured correspondence. This can be done by determining whether there is a target correspondence that is completely consistent with the measured correspondence among multiple pre-stored correspondences, or by determining whether there is a target correspondence that falls within a certain preset error range. This embodiment does not limit this. For example, if the measured current distance is d and the calculated correlation coefficient is c, with a preset error of 5%, the first error range for the distance can be determined to be 0.95d-1.05d, and the second error range for the correlation coefficient can be determined to be 0.95c-1.05c. Thus, among the multiple pre-stored correspondences in the calling device, if there exists a correspondence where the distance is within the first error range and the correlation coefficient is within the second error range, then this correspondence is taken as the target correspondence that matches the measured correspondence, and it is determined that the user of the calling device is wearing a mask.
[0083] By adopting the above technical solution, when the communication device is in a call state, the correlation coefficient between the sound signal from the transmitting device and the sound signal received by the receiving device is determined, and the distance between the transmitting object and the communication device is determined. Since the sound signal received by the receiving device includes the reflected sound signal from the reflecting object, and the distance between the reflecting object and the communication device, as well as whether the reflecting object is wearing a mask, will affect the correlation coefficient between the sound signals, it is possible to determine whether the reflecting object is wearing a mask based on the determined correlation coefficient and the determined distance. This realizes the detection of the user's mask wearing status using the existing sound signal of the communication device, and can achieve one-to-one mask wearing detection for each user of each communication device, without any missed detection. At the same time, it utilizes existing sound resources without the need for additional devices or signals, reducing resource waste.
[0084] The methods provided in the embodiments of this disclosure will be described in detail below.
[0085] In one possible implementation, before determining the first audio signal, i.e., before step S11, this embodiment of the present disclosure may first determine that the calling device is in a handheld calling state. This ensures that the reflected object is the user's face.
[0086] For example, whether the call status of the calling device is a handheld call can be determined by the sensors built into the calling device. For instance, the gyroscope sensor can be used to detect the direction and angle of the calling device during the call status to determine whether it is a handheld call.
[0087] In one possible approach, the second sound signal further includes the direct sound signal from the first sound signal to the receiving device. The mask-wearing detection method provided in this embodiment may also involve first recording the sound signal emitted by the sound-emitting device and the sound signal received by the receiving device to obtain the first sound signal and the second sound signal. Then, the first sound signal recorded within a preset time period is processed into frames to obtain a first sound frame sequence, and the second sound signal is processed into frames to obtain a second sound frame sequence. Finally, the correlation coefficient of each sound frame in the corresponding second sound frame sequence is calculated sequentially based on the first and second sound frame sequences. Determining the correlation coefficient between the second and first sound signals can be achieved by using the second peak value of the calculated correlation coefficient as the correlation coefficient between the second and first sound signals.
[0088] It should be understood that the second sound signal is the sound signal received by the receiver of the communication device, including the direct sound signal of the first sound signal reaching the receiver, and the reflected sound signal of the first sound signal reaching the receiver after being reflected by a reflective object. It may also include ambient sound signals other than the first sound signal and the second sound signal.
[0089] For example, the preset time period can be the time interval from the first moment when the sound signal is emitted by the sound-emitting device of the communication device to the second moment when the sound-receiving device of the communication device receives the reflected sound signal (wherein, the length of this time interval can be determined in advance based on the sound propagation speed and the position of the sound-receiving device of the communication device). Framing the first sound signal within the preset time period can be performed by processing the first sound signal into a first sound frame sequence of length N. For the second sound signal, framing can be performed while recording to obtain a second sound frame sequence of length greater than or equal to N. Then, based on the first and second sound frame sequences, similarity calculations are performed on each sound frame in the corresponding second sound frame sequence to obtain multiple correlation coefficients, each corresponding to the recording time of the sound frame. The second peak value of the calculated correlation coefficient is then used as the correlation coefficient between the second and first sound signals.
[0090] In one possible approach, the correlation coefficient of each sound frame in the second sound frame sequence is calculated based on the first sound frame sequence and the second sound frame sequence. Alternatively, for each sound frame in the second sound frame sequence, a target sound frame sequence is determined, wherein the number of frames in the target sound frame sequence is the same as that in the first sound frame sequence. Then, the correlation coefficient between the target sound frame sequence and the first sound frame sequence is used as the correlation coefficient of the corresponding sound frame.
[0091] For example, the correlation coefficient can be calculated using the following formula:
[0092]
[0093] Where N is the length of the first sound frame sequence and the second sound frame sequence, and C m Let x be the m-th correlation coefficient. n y represents the amplitude of the audio signal corresponding to the nth sound frame in the first sound frame sequence. n+m This represents the amplitude of the audio signal corresponding to the (n+m)th sound frame in the second sound frame sequence.
[0094] If the second sound frame sequence includes the direct sound signal from the first sound signal to the receiving device starting from the k-th sound frame, then the corresponding correlation coefficient C k-1 For the first peak, if the second sound frame sequence, starting from the j-th sound frame, includes the reflected sound signal after the first sound signal has been reflected by the reflecting object, then the corresponding correlation coefficient C is... j-1 This is the second peak. (C) j-1 The correlation coefficient between the second sound signal and the first sound signal was determined.
[0095] It should be understood that the receiver of the communication device is already receiving sound signals when the speaker emits a sound signal. Therefore, before the receiver receives the reflected sound signal after the first sound signal is reflected by the reflective object, it also receives ambient sound signals. After the receiver receives the second sound signal, the speaker continues to emit sound signals. However, the peak time in the second sound signal is later, and the time of the first peak in the first sound signal corresponds to the moment when the speaker first emits a sound signal within a preset time period. Therefore, in the process of calculating the correlation coefficient, for the second sound signal, the same number of sound frames as the first sound frame sequence can be selected as target sound frames in the second sound frame sequence. Then, the correlation coefficient is calculated between the first sound frame sequence and the target sound frame sequence. After each calculation, sound frames in the second sound frame sequence are discarded from front to back, thereby narrowing the calculation range of the second sound frame sequence and determining the correlation coefficient for each corresponding sound frame in the second sound frame sequence.
[0096] For example, if the first audio frame sequence is: x0, x1, x2…x N-1 There are N sound frames in total. When calculating the first correlation coefficient, for each sound frame in the second sound frame sequence, N sound frames can be selected from the beginning. That is, in the second sound frame sequence, starting from the first sound frame, N sound frames are selected as the target sound frame sequence: y0, y1, y2…y N-1 According to the above correlation coefficient calculation formula, the first correlation coefficient C0 is:
[0097]
[0098] At this point, C0 is determined as the correlation coefficient corresponding to the first audio frame.
[0099] When calculating the second correlation coefficient, for each sound frame in the second sound frame sequence, one sound frame is discarded from the beginning, and then N sound frames are selected. That is, in the second sound frame sequence, the first sound frame is discarded, and starting from the second sound frame, N sound frames are selected as the target sound frame sequence: y1, y2, y3…y N According to the above correlation coefficient calculation formula, the second correlation coefficient C1 is:
[0100]
[0101] At this point, C1 is determined as the correlation coefficient corresponding to the second audio frame.
[0102] Of course, other methods can also be used to determine the similarity between the target sound frame sequence and the first sound frame sequence, and this disclosure does not limit this method.
[0103] By using the above method, by determining the same number of sound frames in the second sound frame sequence as the target sound frame sequence, calculating the correlation coefficient between the first sound frame sequence and the target sound frame sequence, and continuously discarding sound frames in the second sound frame sequence from front to back, thereby continuously reducing the number of sound frames in the target sound frame sequence, the correlation coefficient of each sound frame in the second sound frame sequence corresponding to the first sound frame sequence can be determined. This allows for the determination of the peak value and the time of occurrence of the correlation coefficients corresponding to the direct sound signal of the first sound signal reaching the receiving device and the reflected sound signal after the first sound signal is reflected by the reflecting object.
[0104] One possible way to determine the distance between the reflecting object and the communication device is to determine the time difference between the recording times of the sound frames corresponding to the first and second peaks of the calculated correlation coefficient, and then determine the distance between the reflecting object and the communication device based on the time difference and the speed of sound propagation.
[0105] For example, the first peak C can be determined first. k-1 With the second peak C j-1 The corresponding recording times of the sound frames are t k t j Next, determine the speed of sound propagation *v* in the environment where the communication device is located. Then, the distance *d* between the reflecting object and the communication device can be calculated using the following formula:
[0106] d = v × (t) j -t k )
[0107] It should be understood that if the sound-emitting device and the receiver of the communication device are in the same location, the calculated distance d is twice the distance between the sound-emitting device, the receiver, and the reflection point of the reflecting object. If the sound-emitting device and the receiver of the communication device are not in the same location, the calculated distance d is the sum of the distance from the sound-emitting device to the reflection point of the reflecting object and the distance from the reflection point of the reflecting object to the receiver of the communication device. The specific calculation method for the distance between the reflecting object and the communication device can be adaptively adjusted according to the installation position of the sound-emitting device and the receiver of the communication device, the absorption rate of the reflecting object to sound waves, etc., and this disclosure does not limit this. Of course, other methods can also be used to determine the distance between the reflecting object and the communication device, such as using optical distance sensors, infrared distance sensors, etc., and this disclosure does not limit this either.
[0108] Using the above method, based on the measured distance between the reflective object and the call device, it is possible to determine whether there is a matching relationship between the measured distance and the correlation coefficient in the pre-stored correspondence between distance and correlation coefficient, thereby determining whether the user of the call device is wearing a mask.
[0109] In another possible approach, if the calling device does not have a pre-stored target correspondence with distance and correlation coefficient, and it is determined that the user of the calling device is not wearing a mask, then the user of the calling device can be prompted to wear a mask.
[0110] It should be understood that the target mapping already stores the correlation coefficient between the first and second audio signals when a user is wearing a mask and making a call using a handheld device, and the corresponding relationship between this correlation coefficient and the distance between the user's face and the device. Therefore, if no pre-stored correlation is found that matches the measured distance and correlation coefficient in the pre-stored distance-correlation coefficient relationship, it can be determined that the user of the device is not wearing a mask.
[0111] In another possible approach, if it is determined that the user of the calling device is not wearing a mask, the user can be prompted to wear a mask.
[0112] For example, users can be prompted to wear masks by displaying text prompts on the screen of the calling device or by playing voice prompts through the speaking device of the calling device. This embodiment of the disclosure does not limit the method of prompting users of the calling device to wear masks.
[0113] In another possible approach, if it is determined that the user of the calling device is wearing a mask, the sound signal received by the user from the caller's microphone is compensated, and the compensated sound signal is sent to the other party.
[0114] It should be understood that when a user makes a handheld call while wearing a mask, the mask obstructs sound transmission, resulting in sound loss and a lower volume of the user's voice received by the calling device. Therefore, if it is confirmed that the user is wearing a mask, sound compensation can be applied to the sound signal received by the user's microphone on the calling device.
[0115] For example, the pre-stored correspondence between distance and correlation coefficient can be a correspondence between distance, correlation coefficient, and the sound absorption rate of the mask (this correspondence is used to calibrate the mask with this correspondence). Thus, as described above, after determining that there is a target correspondence in the pre-stored correspondence that matches the measured distance and the calculated correlation coefficient, sound compensation can be further performed on the sound signal emitted by the user received by the receiver of the communication device based on the sound absorption rate in the target correspondence.
[0116] For example, the system can also record the user's voice signal when the user is making a handheld call without wearing a mask. If it is determined that the user of the calling device is wearing a mask, the system can compare the voice signal received by the user with the pre-stored voice signal recorded when the user is making a handheld call without wearing a mask. The system can also perform sound compensation on the voice signal received by the user's voice device based on the difference between the audio amplitude of the real-time received voice signal and the audio amplitude of the pre-stored voice signal. The compensated voice signal can then be sent to the other party in the call.
[0117] For example, if it is determined that the user of the calling device is wearing a mask, the real-time received user's voice signal can be amplified by a preset amplification factor, and the amplified voice signal can be sent to the other party in the call. The preset amplification factor can be predetermined based on the absorption rate of the user's voice signal by the mask material. This disclosure does not limit the specific method of voice compensation.
[0118] Figure 4 This is a flowchart illustrating a mask-wearing detection method according to another exemplary embodiment. Figure 4 As shown, this mask-wearing detection method is used in a terminal and includes the following steps:
[0119] S201, It is confirmed that the calling device is in a handheld calling state.
[0120] S202, record the sound signal emitted by the sound-producing device and the sound signal received by the sound-receiving device to obtain the first sound signal and the second sound signal.
[0121] S203, perform frame segmentation processing on the first sound signal recorded within a preset time period to obtain a first sound frame sequence, and perform frame segmentation processing on the second sound signal to obtain a second sound frame sequence.
[0122] S204, for each sound frame in the second sound frame sequence, determine the target sound frame sequence corresponding to that sound frame. The number of frames in the target sound frame sequence is the same as that in the first sound frame sequence.
[0123] S205, the correlation coefficient between the target sound frame sequence and the first sound frame sequence is used as the correlation coefficient of the corresponding sound frame.
[0124] S206, take the second peak of the calculated correlation coefficient as the correlation coefficient between the second sound signal and the first sound signal.
[0125] S207, determine the time difference between the recording times of the sound frames corresponding to the first peak and the second peak of the calculated correlation coefficient.
[0126] S208 determines the distance between the reflecting object and the communication device based on the time difference and the speed of sound propagation.
[0127] S209, determine whether the communication device has a pre-stored target correspondence that matches the distance and correlation coefficient. If the communication device has a pre-stored target correspondence that matches the distance and correlation coefficient, proceed to step S210; otherwise, proceed to step S212.
[0128] S210, confirming that the user of the calling device is wearing a mask.
[0129] S211, perform sound compensation on the sound signal received by the receiver of the communication device from the user, and send the compensated sound signal to the other party in the call.
[0130] S212, It has been determined that the user of the calling device is not wearing a mask.
[0131] S213 prompts users of the calling device to wear masks.
[0132] The above technical solution, when the communication device is in a call, determines the correlation coefficient between the sound signal from the transmitting device and the sound signal received by the receiving device, and determines the distance between the transmitting object and the communication device. Since the sound signal received by the receiving device includes the reflected sound signal from the reflecting object, and the distance between the reflecting object and the communication device, as well as whether the reflecting object is wearing a mask, will affect the correlation coefficient between the sound signals, it is possible to determine whether the reflecting object is wearing a mask based on the determined correlation coefficient and the determined distance. When a user is detected wearing a mask, the user's voice is compensated, and the compensated voice is sent to the other party, thereby improving call quality. When a user is detected not wearing a mask, a timely reminder is given. This achieves the detection of a user's mask-wearing status using the existing sound signal of the communication device, enabling one-to-one mask-wearing detection for each user of each communication device, preventing missed detections. Furthermore, it utilizes existing sound resources without requiring additional devices or signals, reducing resource waste.
[0133] Figure 5 This is a block diagram illustrating a mask-wearing detection device according to an exemplary embodiment. (Refer to...) Figure 5 The device 120 includes a first determining module 121, a second determining module 122, a third determining module 123 and a fourth determining module 124.
[0134] The first determining module 121 is configured to determine, when the calling device is in a calling state, a first sound signal emitted by the speaking device of the calling device and a second sound signal received by the receiving device of the calling device, wherein the second sound signal includes the reflected sound signal of the first sound signal after being reflected by a reflecting object and reaching the receiving device.
[0135] The second determining module 122 is configured to determine the correlation coefficient between the second sound signal and the first sound signal;
[0136] The third determining module 123 is configured to determine the distance between the reflecting object and the communication device;
[0137] The fourth determining module 123 is configured to determine that the user of the calling device is wearing a mask when the calling device has a target correspondence that matches the distance and the correlation coefficient. The target correspondence is the correspondence between the correlation coefficient between the first sound signal and the second sound signal measured when the user is wearing a mask and holding the calling device to make a call, and the distance between the user's face and the calling device.
[0138] Optionally, the second sound signal further includes the direct sound signal from the first sound signal reaching the receiving device, and the device 120 further includes:
[0139] The recording module is configured to record the sound signal emitted by the sound-producing device and the sound signal received by the sound-receiving device, thereby obtaining the first sound signal and the second sound signal;
[0140] The frame segmentation module is configured to perform frame segmentation processing on the first sound signal recorded within a preset time period to obtain a first sound frame sequence, and to perform frame segmentation processing on the second sound signal to obtain a second sound frame sequence.
[0141] The calculation module is configured to calculate the correlation coefficient of each sound frame in the second sound frame sequence in turn based on the first sound frame sequence and the second sound frame sequence.
[0142] The second determining module 122 is configured as follows:
[0143] The second peak of the calculated correlation coefficient is taken as the correlation coefficient between the second sound signal and the first sound signal.
[0144] Optionally, the third determining module 123 is configured as follows:
[0145] Determine the time difference between the recording times of the audio frames corresponding to the first peak and the second peak of the calculated correlation coefficient;
[0146] The distance between the reflecting object and the communication device is determined based on the time difference and the speed of sound propagation.
[0147] Optionally, the computing module is configured as follows:
[0148] For each sound frame in the second sound frame sequence, a target sound frame sequence corresponding to that sound frame is determined, wherein the number of frames in the target sound frame sequence is the same as that in the first sound frame sequence;
[0149] The correlation coefficient between the target sound frame sequence and the first sound frame sequence is used as the correlation coefficient for the corresponding sound frame.
[0150] Optionally, the device 120 further includes a fifth determining module, which is configured to:
[0151] Before determining the first sound signal emitted by the voice output device of the calling device and the second sound signal received by the voice receiver of the calling device, it is determined that the calling device is in a handheld calling state.
[0152] Optionally, the device 120 further includes a sixth determining module, which is configured to:
[0153] If the calling device does not have a pre-stored target correspondence that matches the distance and the correlation coefficient, it is determined that the user of the calling device is not wearing a mask.
[0154] Optionally, the device 120 further includes a prompting module, which is configured to:
[0155] If it is determined that the user of the calling device is not wearing a mask, the user of the calling device will be prompted to wear a mask.
[0156] Optionally, the device 120 further includes a compensation module, which is configured to:
[0157] If it is determined that the user of the calling device is wearing a mask, the sound signal received by the user from the sound receiver of the calling device is compensated, and the compensated sound signal is sent to the other party in the call.
[0158] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0159] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the mask-wearing detection method provided in this disclosure.
[0160] Figure 6 This is a block diagram illustrating an electronic device 800 according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0161] Reference Figure 6 The electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0162] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the mask-wearing detection method described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0163] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0164] Power component 806 provides power to various components of electronic device 800. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0165] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0166] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0167] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0168] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0169] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0170] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the mask-wearing detection method described above.
[0171] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to complete the mask-wearing detection method described above. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0172] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described mask-wearing detection method when executed by the programmable device.
[0173] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0174] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for detecting mask wearing, characterized in that, include: When the communication device is in a call state, a first sound signal emitted by the sound output device of the communication device and a second sound signal received by the sound receiving device of the communication device are determined. The second sound signal includes the direct sound signal of the first sound signal reaching the sound receiving device and the reflected sound signal of the first sound signal reaching the sound receiving device after being reflected by a reflective object. The first sound signal recorded within a preset time period is processed into frames to obtain a first sound frame sequence, and the second sound signal is processed into frames to obtain a second sound frame sequence. Based on the first sound frame sequence and the second sound frame sequence, calculate the correlation coefficient for each sound frame in the second sound frame sequence in sequence. The second peak of the calculated correlation coefficient is taken as the correlation coefficient between the second sound signal and the first sound signal; Determine the distance between the reflective object and the communication device; If the calling device has a pre-stored target correspondence that matches the distance and the correlation coefficient, it is determined that the user of the calling device is wearing a mask. The target correspondence is the correspondence between the correlation coefficient between the first sound signal and the second sound signal measured when the user is wearing a mask and holding the calling device to make a call, and the distance between the user's face and the calling device.
2. The method according to claim 1, characterized in that, Determining the distance between the reflective object and the communication device includes: Determine the time difference between the recording times of the audio frames corresponding to the first peak and the second peak of the calculated correlation coefficient; The distance between the reflecting object and the communication device is determined based on the time difference and the speed of sound propagation.
3. The method according to claim 1, characterized in that, The step of calculating the correlation coefficient for each sound frame in the second sound frame sequence according to the first sound frame sequence and the second sound frame sequence includes: For each sound frame in the second sound frame sequence, a target sound frame sequence corresponding to that sound frame is determined, wherein the number of frames in the target sound frame sequence is the same as that in the first sound frame sequence; The correlation coefficient between the target sound frame sequence and the first sound frame sequence is used as the correlation coefficient for the corresponding sound frame.
4. The method according to any one of claims 1-3, characterized in that, Before determining the first sound signal emitted by the voice-emitting device of the communication device and the second sound signal received by the voice-receiving device of the communication device, the method further includes: It is determined that the calling device is in a handheld calling state.
5. The method according to any one of claims 1-3, characterized in that, Also includes: If the calling device does not have a pre-stored target correspondence that matches the distance and the correlation coefficient, it is determined that the user of the calling device is not wearing a mask.
6. The method according to claim 5, characterized in that, Also includes: If it is determined that the user of the calling device is not wearing a mask, the user of the calling device will be prompted to wear a mask.
7. The method according to any one of claims 1-3, characterized in that, Also includes: If it is determined that the user of the calling device is wearing a mask, the sound signal received by the user from the sound receiver of the calling device is compensated, and the compensated sound signal is sent to the other party in the call.
8. A mask-wearing detection device, characterized in that, include: The first determining module is configured to determine, when the calling device is in a calling state, a first sound signal emitted by the sound output device of the calling device and a second sound signal received by the sound receiving device of the calling device, wherein the second sound signal includes a direct sound signal of the first sound signal reaching the sound receiving device and a reflected sound signal of the first sound signal reaching the sound receiving device after being reflected by a reflecting object. The second determining module is configured to perform frame-segmentation processing on the first sound signal recorded within a preset time period to obtain a first sound frame sequence, and to perform frame-segmentation processing on the second sound signal to obtain a second sound frame sequence. Based on the first sound frame sequence and the second sound frame sequence, calculate the correlation coefficient for each sound frame in the second sound frame sequence in sequence. The second peak of the calculated correlation coefficient is taken as the correlation coefficient between the second sound signal and the first sound signal; The third determining module is configured to determine the distance between the reflecting object and the communication device; The fourth determining module is configured to determine that the user of the calling device is wearing a mask when the calling device has a target correspondence that matches the distance and the correlation coefficient. The target correspondence is the correspondence between the correlation coefficient between the first sound signal and the second sound signal measured when the user is wearing a mask and holding the calling device to make a call, and the distance between the user's face and the calling device.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.
10. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor is configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice data processing method and device, electronic equipment and storage medium
CN113674737A