System and method for detecting a hood sound profile
Patent Information
- Application Number
- CN202311264874.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-05
- Filing Date
- 2023-09-27
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-09-27
AI Technical Summary
然而,这样的方法通常导致在扬声器罩被移除时声音质量下降,并且如果使用不同类型的扬声器罩,则不允许自动重新校准
Smart Images

Figure CN117848486B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure generally relate to the field of acoustic sensing. In particular, this disclosure relates to systems and methods for identifying the acoustic profile of a speaker grille, which can be used to optimize device configuration. Background Technology
[0002] Speaker covers, including speaker grilles, protect speakers and their drivers from external factors such as debris and dust, and also improve the speaker's aesthetic appeal. Without speaker covers, fine dust particles can enter the speaker and degrade sound quality or sound transmission.
[0003] One drawback of using a speaker grille is that it can affect the sound produced. For example, at higher frequencies, the speaker's sound will be affected.
[0004] The conventional method for achieving a speaker grille is to calibrate the speaker to take the grille into account. However, such a method often results in a degraded sound quality when the grille is removed, and automatic recalibration is not allowed if different types of speaker grilles are used. Summary of the Invention
[0005] Therefore, there is a need for systems and methods for identifying the acoustic profile of a speaker grille, enabling the detection of the presence and type of the speaker grille, thereby allowing speaker settings to be optimized to take into account the grille used.
[0006] In one aspect, this disclosure provides a system for detecting a mask sound profile, the mask sound profile comprising at least one sound sample generated based on sound diffracted by the mask. The system includes a device having a housing and at least one speaker driver, at least one microphone, and at least one processor. The device is configured to selectively couple a mask via the housing and generate sound via the at least one speaker driver. The at least one microphone is configured to detect the sound and generate an electrical signal based on the detected sound. The at least one processor is configured to receive the electrical signal, determine using a machine learning algorithm that the spectrogram satisfies or exceeds a similarity threshold of the mask sound profile, and modify at least one characteristic of the device based on adjustment operations associated with the determined mask sound profile.
[0007] In one implementation, at least one sound sample serves as at least one spectrogram, and at least one processor is further configured to convert an electrical signal into a spectrogram. In such an implementation, a cover is determined to be coupled to a housing when the spectrogram satisfies or exceeds a similarity threshold of the cover profile.
[0008] In this implementation, at least one microphone is configured to detect sound only upon receiving a test command. The at least one microphone may be a microphone array.
[0009] The sound may be white noise or generated within the ultrasonic frequency range according to the embodiment.
[0010] In this embodiment, the cover may include one or more of the following: wood, fabric, 3D woven fabric, plastic, marble or other stone, metal such as stainless steel and aluminum, glass, rubber, and leather. The cover may include one or more perforations.
[0011] At least one processor is also configured to update the determined masking sound profile to include a spectrogram according to the implementation.
[0012] In implementation, the machine learning algorithm may combine one or more of the following: image recognition, or mathematical transformations such as the Fast Fourier Transform.
[0013] In one implementation, the determination is based in part on whether a previous spectrogram from a previously detected sound failed to meet or exceed a similarity threshold regarding the masked sound profile.
[0014] In a second aspect, this disclosure provides a method for detecting a mask sound profile, the mask sound profile comprising at least one sound sample generated based on sound diffracted by the mask. The method includes generating sound via a speaker driver, detecting sound via a microphone, generating an electrical signal based on the detected sound, transmitting the electrical signal to a processor, converting the received electrical signal into a spectrogram at the processor, determining, using a machine learning algorithm, that the electrical signal satisfies or exceeds a similarity threshold of the mask sound profile, and altering at least one characteristic of the device based on control operations mapped to the determined mask sound profile.
[0015] In implementation, machine learning algorithms can apply Fast Fourier Transform to electrical signals and / or apply image recognition to the spectrogram of electrical signals.
[0016] The above overview is not intended to describe every embodiment or implementation of the subject matter of this disclosure. The following figures and detailed description illustrate various embodiments in more detail. Attached Figure Description
[0017] The subject matter of this disclosure can be more fully understood by considering the following detailed description of various embodiments in conjunction with the accompanying drawings, in which:
[0018] Figure 1 This is a block diagram of a system for processing the sound profile of a mask, according to an embodiment.
[0019] Figure 2 This is a block diagram of a system for processing the sound profile of a mask, according to an embodiment.
[0020] Figure 3 This is a flowchart of a method for using sound-based touch input according to an implementation method.
[0021] Figure 4A This is a spectrum diagram of a loudspeaker with a fabric cover according to an embodiment.
[0022] Figure 4B This is a spectrum diagram of a loudspeaker with a wooden cover according to an embodiment.
[0023] Figure 5A It is a spectrum diagram of a loudspeaker with a fabric cover according to an embodiment when playing white noise from 15 kHz to 20 kHz.
[0024] Figure 5B This is a spectrum diagram of a loudspeaker with a wooden cover according to an embodiment when playing white noise from 15kHz to 20kHz.
[0025] While various embodiments can be modified into various alterations and alternatives, their details have been shown by way of example in the accompanying drawings and will be described in detail. However, it should be understood that this is not intended to limit the claimed invention to the specific embodiments described. Rather, it is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the subject matter as defined by the claims. Detailed Implementation
[0026] Embodiments of this disclosure relate to systems and methods for identifying when and how sound produced by a loudspeaker is affected by a loudspeaker cover by comparison with known cover sound profiles. Cover sound profiles can be used to identify the presence and type of cover used with a loudspeaker by identifying unique acoustic characteristics associated with each cover. Detection of the cover sound profiles can be used to tune the loudspeaker according to a specific loudspeaker cover to counteract the effect of the loudspeaker cover on the produced sound. The comparison of the produced sound with the cover sound profile can be achieved by applying image recognition to match the spectrogram of the produced sound with the spectrogram included in each cover sound profile. This comparison process can be implemented by applying a machine learning algorithm (MLA) to the image recognition process.
[0027] The inventors of this disclosure have recognized that a speaker grille alters the sound of the speaker, and that this alteration can be measured using a microphone built into the speaker, allowing speaker settings to be corrected to account for different speaker grille effects without any hardware changes to the speaker or grille. In other words, the measurement of the change in sound can be used as a “fingerprint” or grille sound profile for any given speaker grille.
[0028] Embodiments of this disclosure are operable to detect and classify the acoustic profile of a speaker enclosure associated with it without relying on conventional component identification methods such as RFID tags. Therefore, this disclosure is operable for use with existing speakers and speaker enclosures.
[0029] Reference Figure 1 The diagram depicts a block diagram of a system 100 for identifying the presence and type of a speaker grille according to an embodiment. The system 100 can be used to receive and analyze sound generated by a user device 102 and generally includes the user device 102, a network 104, and at least one data source 106.
[0030] User device 102 generally includes a processor 108, a memory 110, at least one transducer 112, and at least one speaker driver 114. Examples of user device 102 include speakers, headphones, earphones, smartphones, tablets, laptops, wearable devices, other consumer electronic devices, or user equipment (UE). For convenience, the term "user device" will be used throughout this document, but does not limit the actual features, characteristics, or composition of any device that may embody user device 102.
[0031] User device 102 may include a housing capable of being removably coupled to cover 116. It is noteworthy that regardless of the design of the housing or cover 116 of user device 102, detection and classification of the cover sound profile associated with a particular speaker cover can be achieved. Therefore, one benefit achieved by embodiments of this disclosure is at least structural and / or material freedom regarding the housing and cover 116 of user device 102.
[0032] Processor 108 can be any programmable device that accepts digital data as input, is configured to process the input according to instructions or algorithms, and provides a result as output. In embodiments, processor 108 can be a central processing unit (CPU), microcontroller, or microprocessor configured to execute instructions of a computer program. Therefore, processor 108 is configured to perform at least basic arithmetic, logic, and input / output operations.
[0033] Memory 110 may include volatile or non-volatile memory, such as that required by the coupled processor 108, to provide not only space for executing instructions or algorithms but also space for storing the instructions themselves. In embodiments, volatile memory may include, for example, random access memory (RAM), dynamic random access memory (DRAM), or static random access memory (SRAM). In embodiments, non-volatile memory may include, for example, read-only memory, flash memory, ferroelectric RAM, hard disk drive, or optical disk drive. The foregoing list is in no way limiting the types of memory that can be used, as these embodiments are given by way of example only and are not intended to limit the scope of this disclosure.
[0034] Transducer 112 refers to any device capable of sensing, detecting, or recording sound to generate an electrical signal, and any device that converts an electrical signal into a sound wave. Transducer 112 can be a cardioid, omnidirectional, or bidirectional microphone. In embodiments, transducer 112 can be a single microphone or a microphone array comprising multiple microphones. Multiple microphones can be used to distinguish subtle differences in detected sound and determine the angle of arrival of the sound. In some embodiments, transducer 112 can be a piezoelectric transducer. In other embodiments, transducer 112 can be combined with other types of acoustic sensors or combinations of sensors or devices that together can sense sound, pressure, or other characteristics related to audible or inaudible sounds generated by contact with a surface (relative to the sensitivity of human hearing). Such inaudible sounds can include ultrasound. Transducer 112 can be configured to record and store digital sound or data obtained from captured sound. Any signals generated by transducer 112 can be sent to processor 108 for analysis.
[0035] In one embodiment, at least one transducer 112 can convert electrical signals into sound, such as a speaker driver. In such an embodiment, the transducer 112 may be a speaker driver configured to produce a specific portion of the audible frequency range, such as a super tweeter, tweeter, midrange driver, woofer, subwoofer, and rotary woofer. In another embodiment, one or more speaker drivers may be incorporated into the user device 102.
[0036] Implementations of system 100 generally include at least one transducer for generating sound and at least one transducer for receiving sound. Although user device 102 is described as a single device, it should be understood that the functionality and components of the user device may be separated among one or more devices. For example, a microphone may be located outside user device 102 and detect sound generated by the user device, such that the microphone can determine whether a speaker grille is present on user device 102.
[0037] System 100 can be implemented regardless of the number or type of transducers 112, although in some embodiments, it may be advantageous to arrange one or more transducers 112 in known locations relative to the housing. In embodiments, the transducers 112 may be located inside or outside the housing or stored in a housing independent of the user device. For example, the transducer 112 may be positioned inside a telephone and configured to estimate the cover sound profile from an external speaker. The position of the transducer 112 relative to the housing allows for more accurate cover sound profiles of individual devices because perceptible differences in sound from device arrangement can be mitigated or otherwise reduced. In embodiments, the transducer 112 may be configured to detect sound frequencies in the range of 1 Hz to 80 kHz. In embodiments, the transducer 112 may be configured to detect sound frequencies in the range of 1 kHz to 20 kHz. In embodiments, the transducer 112 may be configured to detect ultrasonic frequencies in the range of 19 kHz to 22 kHz.
[0038] The speaker grille 114 can be any speaker grille configured for use with the user device 102. In one embodiment, the grille 114 can be removably coupled to the user device 102, allowing the user to customize their speaker by selecting speaker grilles with different designs, materials, and colors. In another embodiment, the grille 114 can include one or more of the following: wood, fabric, plastic, marble, metal, glass, rubber, and leather. For example, the grille 114 can include a molded plastic grille with perforations to allow sound to pass through, which is then covered with fabric. In other embodiments, the grille 114 can be made entirely of a material with perforations, such as wood. Typically, the perforations in the grille 114 have a significant impact on the overall sound profile of the grille. Therefore, embodiments of this disclosure are particularly effective when the perforations in each grille are located in unique positions or have varying dimensions.
[0039] One arrangement that can be detected by this disclosure is the absence of a speaker grille. A speaker without a grille can be characterized by different grille sound profiles. In the absence of a grille, it can be recommended to the user that a grille be used to extend the speaker's lifespan.
[0040] Embodiments of this disclosure can be used with any material coupled to or removably coupled to a speaker. For example, a sound profile can be used with headphones to detect differences in headphone pads. Using headphones as an example, the sound produced by the headphones can be adjusted according to the material used for the headphone pads.
[0041] User device 102 may include other features, devices, and subsystems, such as an input / output engine or a sound processing engine, which include various engines or tools, each of which is constructed, programmed, configured, or otherwise adapted to autonomously perform a function or a set of functions. As used herein, the term "engine" is defined as a physical device, component, or arrangement of components implemented using hardware—for example, via an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)—or a combination of hardware and software—such as a microprocessor system and a set of program instructions that adapt the engine to perform a specific function, which (when executed) translates the microprocessor system into a dedicated device. An engine may also be implemented as a combination of both, where a specific function is implemented solely by hardware, while other functions are implemented using a combination of hardware and software. In certain implementations, at least a portion of the engine, and in some cases, all of the engine, may execute on one or more processors of one or more computing platforms consisting of hardware that executes operating systems, system programs, and applications (e.g., one or more processors, data storage devices such as memory or drive memory, input / output devices such as network interface devices, video devices, keyboards, mice, or touchscreen devices, etc.), while also implementing the engine using multitasking, multithreading, appropriate distributed (e.g., cluster, peer-to-peer, cloud, etc.) processing or other such technologies. Therefore, some or all of the functions of processor 108 may execute in various physically feasible configurations within each engine and should not be limited to any particular implementation illustrated herein unless such limitation is explicitly stated.
[0042] User device 102 is configured to provide bidirectional data communication with network 104 via a wired or wireless connection. The specific design and implementation of the input / output engine of processor 108 may depend on the communication network in which user device 102 is intended to operate. User device 102 can access stored data from at least one data source 106 via network 104.
[0043] Data source 106 may be a general-purpose database management storage system (DBMS) or relational DBMS implemented using solutions such as Oracle, IBM DB2, Microsoft SQL Server, PostgreSQL, MySQL, SQLite, Linux, or Unix, trained to interpret spectrograms or sound samples corresponding to the speaker grille contours. Data source 106 may store one or more training datasets configured to facilitate future image recognition of speaker grille contours within the spectrograms of captured sounds, or to otherwise identify speaker grille contours by analyzing captured sounds. In embodiments, such analysis may include applying mathematical transformations (e.g., Fourier transform, fast Fourier transform, wavelet transform) to better identify the acoustic characteristics associated with the speaker grille. In embodiments, data source 106 may classify or implement the training dataset based on acoustic characteristics of the generated sounds detected, such as high-frequency reduction. In embodiments, data source 106 may be a native data source of user device 102, eliminating the need for a connection to network 104.
[0044] One purpose of data source 106 is to store multiple spectrograms, which are visual representations of the signal strength or "loudness" of a signal at various frequencies present in a specific waveform over time. Spectrograms provide a visual representation of the presence of more or less energy and how the energy level changes over time. These visual representations can be an effective way to compare and analyze detected sounds. Spectrograms can be depicted as heatmaps, that is, as images that show intensity by changing color or brightness.
[0045] A spectrogram can be generated based on known sounds produced by a known enclosure. Each enclosure can continuously alter the produced sound. These altered sounds can then be converted into a spectrogram and stored within an enclosure sound profile associated with a specific enclosure. Differences in the detected sounds within a specific enclosure sound profile, such as those caused by the distance between the receiving transducer and the enclosure, can be learned across various robust sample sizes, where each sample represents the sound produced using the corresponding installed speaker enclosure.
[0046] Raw frequency data can also be used to detect sound characteristics without generating a spectrogram. When using raw frequency data, isolating the frequencies of interest can improve accuracy. For each speaker enclosure sound profile, a range of relevant frequencies can be identified. This frequency range can represent the frequencies that are most different from other speaker enclosure sound profiles (or the sound produced without speaker enclosures). For example, if a speaker enclosure arrangement produces a high energy level between 18 kHz and 18.5 kHz, the analysis can be limited to this range, as well as any different range of relevant frequencies used for other speaker enclosure sound profiles, to reduce the amount of comparison data and thus simplify calculations.
[0047] The inventors of this disclosure have recognized that, over time, speaker enclosures generally have a similar effect on the energy levels of the produced sound. For example, a particular type of wooden enclosure typically produces sound with common visual characteristics (e.g., reduced high frequencies) when converted to a spectrogram, as can be achieved by... Figure 4B and Figure 5B The spectrum of the wood cover shown in the figure is similar to that shown in the figure. Figure 4A and Figure 5A The spectrograms of the fabric cover shown are compared to those in the diagram. In this implementation, artificial intelligence (AI) or MLA can be trained to identify the spectrograms associated with the speaker cover by performing image recognition on these visual characteristics.
[0048] The visual characteristics of the spectrogram associated with a particular speaker grille can be extracted manually or automatically using machine learning methods, such as neural networks, to generate a spectrogram of the grille's acoustic profile. These grille acoustic profiles, each including a set of visual characteristics associated with a particular speaker grille, can then be stored in data memory 106 for future image recognition comparison with spectrograms of detected audio. Such comparisons can be made by calculating a similarity metric using correlation or machine learning regression algorithms. For example, if the similarity between the spectrogram and the grille acoustic profile is higher than a certain threshold (e.g., at least 75%, at least 90%, at least 95%, or at least 99%), the matching process can determine that the spectrogram represents the grille associated with the grille acoustic profile.
[0049] In implementation, MLA can extract visual characteristics from specific portions of the spectrogram to better compare the effect of the cover on the generated sound. In some cases, segmenting or enhancing specific regions of the spectrogram during image recognition analysis can improve accuracy by limiting the influence of outliers (e.g., objects blocking part of the speaker cover).
[0050] MLA technology can be applied to both labeled (supervised) and unlabeled (unsupervised) spectrogram data. Furthermore, the classifier can accept parameters such as device type (e.g., speakers and headphones may have different parameters or sound detection capabilities). Reasons for employing such a classifier include identifying the position of the transducer relative to the user device's housing or the number of transducers present.
[0051] In operation, the spectrogram can be processed by MLA to benefit from background information of the detected sound. In other words, MLA allows for processing of variable-length inputs and outputs by maintaining state information over time. In one example, if no masked sound profile was detected in a previous attempt, the speaker can repeat white noise for a certain period of time. MLA can take the previous attempts into account and adjust the subsequent analysis of the second spectrogram accordingly, such as lowering the threshold required to consider a specific masked sound profile match that was the closest match in previous comparisons. Similarly, if multiple spectrograms are generated within a short time window, commonalities, such as enhanced background noise, can be identified and interpreted. Therefore, the environment surrounding the detected sound can help AI gain insights that take into account environmental changes.
[0052] In implementations, training data may include multiple spectrograms based on known sounds, acoustic impulses, or white noise. These sounds can then be played on a speaker, and the resulting spectrograms can be compared to masked sound profiles. In some implementations, the MLA can be trained to recognize masked sound profiles during normal playback through a speaker. This can be achieved if the generated spectrograms represent frequently used media, such as songs, or if standard operating sounds (e.g., power-on noise or connection establishment alarms) are processed as the basis for the training set. With sufficient training based on these examples, the MLA can better identify when frequencies observed in the spectrograms represent masked sound profiles in standard device usage. This analysis can be improved during operation by including feedback loops for common device triggering conditions.
[0053] MLA can be trained by generating white noise or ultrasonic frequencies for a specific time period for a desired device according to an embodiment. Since sounds inaudible to the human ear can be used, the detection and identification of speaker enclosures can occur during device use without interrupting the user experience. Spectrograms of these known sounds can be used to improve the recognition of the enclosure's sound profile, while unknown sounds may require, for example, more detailed spectrogram comparisons over longer time periods to provide reliable detection of known enclosures. Similarly, spectrograms from known devices can improve the speed and accuracy of image recognition comparisons adapted to variations that may arise due to differences in device configuration.
[0054] In implementations, MLA can be simplified to reduce computational and power requirements. This simplification can be achieved by reducing the length of the recorded sound samples or by analyzing sound samples in the frequency domain to determine the presence and / or intensity of key frequencies. For example, in an implementation that does not generate a spectrogram for image recognition analysis, the power of the frequencies of a single sound sample recorded over a time period (e.g., 50 milliseconds or 3 seconds) is analyzed to determine the most likely speaker grille being used. This analysis may include a lookup table or a mapping of frequencies attenuated by a known speaker grille. In cases of limited computational power, a "best guess" method can be used, which relies on a lower similarity threshold and reduced sound sample duration to simplify grille matching to a lookup operation. In implementations, simplified MLA can run on devices with limited computational power, such as wearable devices. In such implementations, the lookup table or mapping can be performed by applying MLA to sound samples (training data) without converting to a spectrogram.
[0055] In implementations that rely on frequency comparison, the time component of the sound samples does not need to be considered. Since speaker grilles typically cause sound to decay continuously over time, a purely frequency-based value comparison can be used, rather than the frequency-over-time comparison used when performing image recognition of a spectrogram. This frequency approach is particularly advantageous when the speaker is configured to generate white noise or a set range of ultrasonic frequencies to cover the acoustic profile during the comparison process.
[0056] Reference Figure 2 The diagram depicts a block diagram of a user device 102 according to an embodiment. As depicted, the user device 102 may include a transducer 112 and a speaker driver 114. Figure 2 Possible paths of sound 118 generated by speaker driver 114 and received by transducer 112 are also shown. The effect of the cover 116 can be determined by comparing the spectrogram generated by sound 118 with a spectrogram associated with a known cover for user device 102. In embodiments where sound 118 is used to establish the cover sound profile, identification can be more accurate.
[0057] Reference Figure 3 The document describes a method 300 for detecting the presence and type of a speaker enclosure using sound, according to an embodiment. Method 300 can be implemented via a user device, such as user device 102.
[0058] At point 302, a test command may optionally be received by the user device to prompt the user device to check for an updated cover. The test command can be communicated via the user device's UI or prompted by different audio cues. For example, the user device may passively listen to different audio cues, such as spoken phrases, which may then prompt the user device to generate an sound to test for the presence of a cover. The arrangement of the test command can extend the user device's battery life, possibly due to the continuous or periodic processing of sound to determine the presence or alteration of a cover. In some embodiments, any user interaction with the user device or devices associated with the user device can be used as a test command. Once the test command has been received by the user device, a prompt or alert can be delivered to the user to indicate that a test cycle has begun. The duration of this test cycle can be customized based on user preferences or user device considerations. In some embodiments, operational sounds, such as power-on sounds, can be used as test commands.
[0059] At point 304, the generated sound is detected by the transducer of the user equipment. In this implementation, the time period for sound detection can be shortened or lengthened based on the known mask sound profile and its relative uniqueness compared to other masks for a particular user equipment. For example, if the mask sound profile continuously blocks sound of a specific frequency that is otherwise detected by the user equipment, the user equipment can stop listening after a relatively short period of time during which the specific frequency should have been received but was not, thus indicating the presence of a known mask blocking such frequencies. Therefore, the mask sound profile can be detected more effectively where the specific frequency is omitted or amplified.
[0060] In an implementation that does not rely on test commands, the user device can passively listen to sound at point 304. Sound interpretation technology allows the user device to selectively process detected sound. For example, the user device can process only sound that has been calculated as having been generated by the user device. Parameters that can be effectively used as test commands may include sound volume, sound direction, estimated sound location, sound characteristics, etc.
[0061] According to one implementation, upon detecting a sound profile, a confirmation alarm or audible pulse can be optionally presented to the user to prevent unnecessary device adjustments. In such an implementation, the user can then confirm the desired adjustment via a repeat test command, user input associated with a confirmation operation, voice command, or the like.
[0062] In some implementations, sound profile mapping can be used to adjust user device settings in real time. In one example, a user can play one or more sounds during standard operation of the user device. These sounds can then trigger the device to begin continuous image recognition of the detected sounds to cover the sound profile. Such real-time analysis can facilitate acute control over device settings and user preferences. In implementations, real-time sound analysis can be simulated through frequent, periodic sampling of sound output.
[0063] For user devices incorporating sensors, sensor data can be used to determine whether a check should be performed for a new cover. For example, if a speaker includes a proximity sensor, a proximity reading can trigger a cover check because a user replacing the cover might result in such a proximity reading.
[0064] At point 306, the transducer converts the detected sound into an electrical signal, which is then sent to the sound processing engine. In this embodiment, the sound processing engine may be a processor of the user device or a server located external to the user device, such as one coupled to the user device via a network communication.
[0065] At point 308, the electrical signal is converted into a spectrogram by the sound processing engine. In an implementation, the resulting spectrogram can be processed to enhance the distinguishability of the masked sound profile. For example, specific audio tones or frequencies can be shifted to simplify the comparison process with trained masked sound profiles, or otherwise improve matching accuracy.
[0066] At 310, an image recognition MLA is applied to the resulting spectrogram to determine if the mask sound profile matches at 312. This comparison is achieved using an image recognition MLA trained on a dataset of mask sound profiles associated with known masks. In this implementation, the MLA can be trained on supervised data, such that the trained spectrogram is labeled with associated masks. By labeling the MLA's training data, mask sound profiles sharing common characteristics (e.g., perforation location or material) can be associated to improve image recognition accuracy. Notably, over time, the MLA can establish similar relationships between mask sound profiles using unsupervised data, but such training may be less efficient.
[0067] At point 314, if a masking sound profile is detected, an adjustment operation for the user device is performed. The masking sound profile can be mapped to different adjustment operations for the user device and can be based on a single user preference or profile. The adjustment operation can be one or more of the following: raising or lowering a specific frequency or frequency range, changing dynamics or beamforming characteristics, volume control, and changing the operating mode of the user device. In an embodiment, the adjustment operation can be used for a user device separate from the device detecting the audio.
[0068] At 316, improvements may optionally be implemented to enhance future processing of sound detection. These improvements may include one or more feedback loops designed to improve future masking contour recognition or personalization of user adjustment actions.
[0069] In implementation, improvements can be based on the background of the detected sound. For example, if the spectrogram generated following the test command does not match the mask sound profile, a temporary flag can be proposed to indicate the most recent failed match. In subsequent iterations of method 300, if another known sound is detected after a subsequent test command and a failed match flag exists, the threshold required for the known sound to be considered a match with the mask sound profile can be changed. Such an arrangement can reduce user frustration when the user attempts to identify changes in the mask but repeatedly fails.
[0070] In implementations, one or more feedback loops can modify parameters of the image recognition MLA. These parameters may include one or more of the following: the duration of a failed match flag in each case; the strength of the matching threshold that changes in response to the proposed failed match flag; whether the matching threshold changes universally or only for one or more mask voice contours identified as most similar to previous failed attempts; and issuing an alert or notification to the user that a mask voice contour has not been recognized. In some implementations, the impact of failed mask voice contour matching can be amplified during failed attempts or configured to be enabled once a certain number of failed attempts have been made. In other words, once an attempt fails, the similarity threshold can be lowered so that adjustments can be made to match the closest mask voice contour, even when using an unrecognized mask.
[0071] It should be understood that the various operations used in the methods of this teaching can be performed in any order and / or simultaneously, as long as the teaching remains operational. Furthermore, it should be understood that the apparatus and methods of this teaching may include any number or all of the embodiments described in the embodiments, as long as the teaching remains operational.
[0072] As previously mentioned, the mask sound profile (i.e., the spectrogram associated with a particular mask) can be mapped to different device tuning operations. In operation, a set of potential tunings can be presented to the user, and then the desired tuning operation can be customized based on the detected mask. These user preferences can then be associated with a user profile, enabling tuning to be applied to any device interacting with the user profile.
[0073] Reference Figure 4A , Figure 4B , Figure 5A and Figure 5BThe spectrum diagrams of different covers according to different implementation methods are shown. Figure 4A This is a spectrum diagram showing the sound profile of the fabric front cover of a loudspeaker. Figure 4B It shows the relationship with Figure 4A A spectrum diagram of the sound profile of the wooden front cover of the same speaker. Figure 5A This is a spectrum diagram showing the white noise from 15kHz to 20kHz generated by a speaker with a front fabric cover. Figure 5B It shows the relationship between and Figure 5A A spectrum diagram of white noise from 15kHz to 20kHz produced by the same speaker but with a front wooden cover.
[0074] Differentiating the type of speaker grille installed can improve the sound of user devices while simplifying and streamlining the user experience. Acoustic testing of speaker grille types allows for the identification of grilles used with existing speakers and their easy updating to add a grille sound profile for new grilles.
[0075] Therefore, embodiments of this disclosure can improve the sound of existing user devices without requiring hardware modifications or user interaction. If another device can detect the sound, even devices without microphones can be controlled or manipulated using the disclosed methods. Thus, compared to cover detection methods involving incorporating active sensors or RFID tags into the cover, cover sound profile recognition can significantly reduce production costs. Furthermore, the cover sound profile recognition of this disclosure allows users to obtain a better listening experience without manual interaction with the speaker or speaker cover.
[0076] Various embodiments of the systems, apparatus, and methods have been described herein. These embodiments are given by way of example only and are not intended to limit the scope of the claimed invention. Furthermore, it should be understood that the various features of the described embodiments can be combined in various ways to produce many additional embodiments. Moreover, although various materials, sizes, shapes, configurations, and positions, etc., for the disclosed embodiments have been described, other materials, sizes, shapes, configurations, and positions, etc., besides those disclosed, may be used without departing from the scope of the claimed invention.
[0077] Those skilled in the art will recognize that the subject matter of this disclosure may include fewer features than those described in any of the individual embodiments above. The embodiments described herein are not intended to be an exhaustive representation of how the various features of the subject matter of this disclosure can be combined. Therefore, embodiments are not mutually exclusive combinations of features; rather, as will be understood by those skilled in the art, various embodiments may include combinations of different individual features selected from different individual embodiments. Furthermore, unless otherwise stated, elements described with respect to one embodiment may be implemented in other embodiments, even if not described in those embodiments.
[0078] While dependent claims may refer to a specific combination with one or more other claims in the claims, other embodiments may also include combinations of dependent claims with the subject matter of each other dependent claim, or combinations of one or more features with other dependent or independent claims. Such combinations are set forth herein unless stated otherwise.
[0079] Any combination of references to the foregoing documents is limited to such that no subject matter contrary to the express disclosure herein is incorporated. Any combination of references to the foregoing documents is further limited to such that any claims included in the documents are not incorporated herein by reference. Any combination of references to the foregoing documents is further limited to such that any definitions provided in the documents are not incorporated herein by reference unless expressly included herein.
[0080] For the purposes of interpreting the claims, unless the claims refer to the specific terms “method for…” or “step for…”, it is expressly indicated that the provisions of 35 U.SC §112(f) do not apply.
Claims
1. A system for detecting a mask sound profile, the mask sound profile comprising at least one sound sample generated based on sound diffracted by the mask, the system comprising: The device includes a housing and at least one speaker driver, the device being configured to: The cover is selectively coupled via the housing; Sound is generated via the at least one speaker driver; At least one microphone, said at least one microphone being configured to: Detect the sound; and An electrical signal is generated based on the detected sound; and At least one processor, said at least one processor being configured to: Receive the electrical signal; Machine learning algorithms are used to determine whether the electrical signal meets or exceeds a similarity threshold for the mask sound profile; and At least one characteristic of the device is altered based on adjustment operations associated with the determined sound profile of the mask.
2. The system according to claim 1, wherein, The at least one sound sample serves as at least one spectrogram, and the at least one processor is further configured to convert the electrical signal into a spectrogram prior to determination.
3. The system according to claim 1, wherein, The at least one microphone is configured to detect the sound only when a test command is received.
4. The system according to claim 1, wherein, The sound is white noise.
5. The system according to claim 1, wherein, The sound is generated at ultrasonic frequencies.
6. The system according to claim 1, wherein, The mask sound profile is determined from a plurality of mask sound profiles corresponding to different types of masks, wherein the different types of masks include masks composed of one or more of wood, fabric, plastic, marble, metal, glass, rubber, and leather, and wherein the adjustment operation is based on the mask type associated with the determined mask sound profile.
7. The system according to claim 1, wherein, The at least one microphone is a microphone array.
8. The system according to claim 1, wherein, The machine learning algorithm incorporates Fast Fourier Transform.
9. The system according to claim 1, wherein, The machine learning algorithm is combined with image recognition.
10. The system according to claim 1, wherein, The at least one processor is also configured to update the determined mask sound profile to include the electrical signal.
11. The system according to claim 1, wherein, The determination is partly based on whether a previous electrical signal from a previously detected sound failed to meet or exceed the similarity threshold of the masked sound profile.
12. A method for detecting a mask sound profile, the mask sound profile comprising at least one sound sample generated based on sound diffracted by the mask, the method comprising: Sound is generated via a speaker driver; The sound is detected via a microphone; An electrical signal is generated based on the detected sound; The electrical signal is transmitted to the processor; The recognition machine learning algorithm is used to determine whether the electrical signal meets or exceeds the similarity threshold of the sound profile of the mask; as well as At least one characteristic of the device is changed based on control operations mapped to the determined acoustic profile of the mask.
13. The method according to claim 12, wherein, The at least one sound sample serves as at least one spectrogram, and the method further includes converting the received electrical signal into a spectrogram at the processor prior to determination.
14. The method of claim 12 further includes receiving a test command before detecting the sound.
15. The method according to claim 12, wherein, The sound is white noise.
16. The method according to claim 12, wherein, The sound is generated at ultrasonic frequencies.
17. The method of claim 12, wherein the cover comprises one or more of the following: wood, fabric, plastic, marble, metal, glass, rubber, and leather.
18. The method according to claim 12, wherein, The at least one processor is also configured to update the determined mask sound profile to include the electrical signal.
19. The method according to claim 12, wherein, The machine learning algorithm is combined with image recognition.
20. The method according to claim 19, wherein, The determination is partly based on whether a previous electrical signal from a previously detected sound failed to meet or exceed the similarity threshold of the masked sound profile.
Citation Information
Patent Citations
Electronic device and electronic device control method
US20190333473A1
Method and system for vision-based defect detection
US20210116293A1