Apparatus and methods for use with treatment devices

JP2025506503A5Pending Publication Date: 2025-11-17KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024547654
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-15
Filing Date
2023-01-30
Publication Date
2025-11-17

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus and method for use with a treatment device configured to perform a treatment action on a body part of a subject, the treatment device configured to apply light pulses to the skin of the body part to perform the treatment action and to output a flash sound for each light pulse. In one embodiment, a trained algorithm or computing system is used to reliably detect the flash sound and notify the user to avoid over-treatment.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an apparatus and method adapted for use with a treatment device adapted to perform a treatment treatment on a body part of a subject, the treatment device adapted to apply light pulses to the skin of the body part to perform the treatment treatment, and to generate a sound flash for each light pulse generated by the treatment device. The present invention further relates to a computer-implemented method of such an apparatus, a system comprising the treatment device, and a computer program. [Background technology]

[0002] Techniques for removing unwanted hair include shaving, electrolysis, plucking, laser and light therapy (known as photoepilation), and injections of therapeutic antiandrogens. Light-based techniques are also used for other types of skin treatments, such as inhibiting hair growth and treating acne.

[0003] By using the appropriate configuration of light energy (i.e., in terms of its wavelength, intensity, and / or pulse duration (if the light is pulsed)), selective heating of the hair root and subsequent temporary or permanent damage to the hair follicle can be achieved. Home photoepilation devices, such as the Philips Lumea device, use intense pulsed light (IPL) from a high intensity light source, such as a xenon flash lamp that produces high power bursts of broad spectrum light.

[0004] Photoepilation treatments are characterized in that a user of photoepilation treats a relatively small area of ​​skin for the purpose of hair removal. Photoepilation uses intense light to heat the melanin in the hair and hair roots, which causes the hair follicles to go dormant and prevent hair regrowth. To effectively use this technology for hair removal, the user must completely treat the skin without leaving any gaps. The effect only lasts for a limited period of time, so treatments must be repeated periodically. Typically, treatments are performed once every two weeks for an initial period of about two months, followed by once every four to eight weeks for a maintenance period.

[0005] The Lumea device is supported by a Lumeaapp (i.e., an application program or software application runnable on a computer, laptop, tablet, and / or smartphone that can be wirelessly coupled, for example, to the treatment device) to assist a user in the hair removal process, thereby improving the user experience and engagement of the Lumea device.

[0006] In many photoepilation devices, each flash / burst of light emitted for treatment is immediately followed by a sound flash, which is a short impulse of sound that is specific in both time and frequency (e.g., about 1 second in time and about 9 kHz in frequency). The number of flashes required for treatment varies depending on the size of the body part to be pulsed. Sustained exposure to high-intensity laser pulses for a time longer than recommended may cause overtreatment and serious damage to the skin.

[0007] WO 2019 / 079972A1 discloses a specific voice recognition method that can recognize a specific voice with a simple algorithm, low computational complexity, and low hardware equipment requirements. The method includes the steps of sampling a voice signal, obtaining a Mel-frequency cepstral coefficient characteristic parameter matrix of the voice signal, and extracting characteristic parameters from the characteristic parameter matrix. The characteristic parameters are provided to a deep neural network-based specific voice feature model to determine whether the voice signal has the specific voice feature. The specific voice recognition device is configured to receive a user's voice and determine whether the voice has the specific voice feature. An example of the specific voice feature mentioned in WO 2019 / 079972A1 is a coughing sound. Summary of the Invention [Problem to be solved by the invention]

[0008] It is an object of the present invention to provide a means to avoid or at least reduce the risk of overtreatment. [Means for solving the problem]

[0009] In a first aspect of the invention, there is provided an apparatus configured for use with a treatment device configured to perform a treatment treatment on a body part of a subject, the treatment device configured to apply light pulses to skin of the body part to perform the treatment treatment, and to generate a flash sound for each light pulse generated by the treatment device, the apparatus comprising: an input configured to obtain an audio signal sensed during a treatment procedure; A processing unit and an output unit, The processing unit is Segmenting the audio signal into audio segments, in particular audio segments of a given or pre-determined duration; Extracting one or more characteristic features for each audio segment from the audio segments; detect, for each audio segment, whether the audio segment includes a flash sound generated by the treatment device for each light pulse by applying a trained algorithm or computing system that has been trained with a plurality of audio segments, each of the audio segments including one or more of the characteristic features, to detect within the audio segment a flash sound generated by the treatment device for each light pulse; configured to vary a flash sound counter that counts the number of flash sounds detected since the start of a treatment procedure or a predetermined or user-instructed moment; The output section is Flash sound counter, an indication of whether the flash sound counter has exceeded the flash sound threshold; and Indication that the treatment procedure is complete An apparatus is presented that outputs one or more of:

[0010] In a second aspect of the invention, there is provided an apparatus configured for use with a treatment device configured to perform a treatment procedure on a body part of a subject, the treatment device configured to apply light pulses to skin of the body part to perform the treatment procedure and to generate a flash sound for each light pulse generated by the treatment device, the apparatus comprising: an input for receiving an audio signal sensed or artificially produced during a treatment procedure; a processing unit, The processing unit is Segmenting the audio signal into audio segments, in particular audio segments of a given or pre-determined duration; Extracting one or more characteristic features for each audio segment from the audio segments; deriving, for each audio segment, flash sound information from the annotation of the audio signal indicating whether the audio segment includes a flash sound generated by the treatment device for each light pulse, the annotation of the audio signal indicating within the audio signal the flash sound generated by the treatment device for each light pulse, in particular time information of the flash sound generated by the treatment device for each light pulse; An apparatus is presented that is configured to provide the extracted one or more characteristic features and derived flash sound information to an algorithm or computing system to train the algorithm or computing system to detect, for each audio segment, a flash sound generated by a treatment device for each light pulse within the audio segment based on the one or more characteristic features.

[0011] In an embodiment of the invention, an apparatus comprises: displaying a first screen requesting user input regarding a treatment procedure; acquiring user input via a first screen, the user input including an image of a body part of the subject to be treated and / or size information of a size of the body part of the subject to be treated; Displaying a second screen indicating connectivity between the device or an electronic device comprising the device and a treatment device; the operating state of the skin treatment device; a flash sound counter for counting the number of flash sounds detected since the start of a treatment procedure or a predetermined or user-initiated moment; an indication of whether the flash sound counter has exceeded the flash sound threshold; and Indication that the treatment procedure is complete and an output unit configured to control the user interface to display a third screen indicating at least one of the following:

[0012] In a third aspect of the present invention, a treatment device comprising one or more light sources generating light pulses for application to skin of a subject's body to perform a treatment procedure, the treatment device being configured to generate a flash sound for each light pulse generated by the treatment device; an audio sensor configured to sense an audio signal during a treatment procedure; A system is presented, comprising an apparatus according to any one of the various aspects disclosed herein.

[0013] In another further aspect of the invention, there is provided a corresponding (computer-implemented) method, a computer program comprising program code means for causing a computer to perform the steps of the respective method disclosed herein when said computer program is run on a computer, as well as a non-transitory computer readable recording medium storing said computer program product which, when run by a processor, causes a respective method disclosed herein to be performed.

[0014] Preferred embodiments of the invention are defined in the dependent claims. It is to be understood that the claimed methods, systems, computer programs and media have similar and / or identical preferred embodiments as the claimed apparatus as specifically defined in the dependent claims and disclosed herein.

[0015] All aspects of the present invention are based on a common idea that, considering the harmful aspects of laser hair removal treatment, the user should know exactly how many flashes are required for the treatment and / or know when the desired area has been treated with a sufficient number of flashes. Typically, a single treatment requires, for example, about 30-40 flashes (depending on the length of the desired treatment area). The user can manually count the number of flashes applied during one treatment session, but there is no recommendation for any number of flashes, which leads to a less than satisfactory user experience. On the other hand, the user may be otherwise preoccupied with other tasks or distracted by background noise, which may make it difficult to accurately track the total number of flashes. Therefore, it is provided by the present invention (first aspect) to use a trained algorithm or computing system to automate the counting of flash sounds (and therefore the number of flashes emitted) during a treatment session. Such a trained algorithm or computing system can be trained to accurately count the total number of flashes during each treatment and to deal with background noise generated by both the treatment device and the environment. Furthermore, the non-periodic nature of the signal is taken into account, thus improving the accuracy in obtaining the number of flashes emitted by the skin treatment device during a treatment session.

[0016] In the context of the present invention, varying a counter for counting the number of flash sounds during treatment is to be understood as such that each time a flash sound is detected the counter value is incremented (e.g. starting from zero) or decremented (e.g. starting from the estimated maximum counter value for treating a treatment area).

[0017] Training, according to the present invention (second aspect), is performed using a number of audio segments, each of which includes one or more characteristic features derived from the audio segment (e.g., segments of, for example, 1 second duration), by detecting flash sounds within the audio segments.

[0018] Further, according to the present invention (third aspect), a user interface is provided that efficiently captures and presents various information to the user, including the number of flashes emitted during a treatment session, which helps the user easily understand the current progress of the treatment and avoid over-treatment.

[0019] Preferably, the present invention uses a learning system, machine learning, classification model, binary classification model, neural network, or convolutional neural network as a trained algorithm or computing system. For example, in one implementation, an XGBoost classifier as known in the art may be used. XGBoost is a decision tree-based ensemble machine learning algorithm that uses a gradient boosting framework. It can be used in a wide range of applications, such as to solve regression, classification, ranking, and user-defined prediction problems. XGBoost provides high prediction performance with minimal processing time, both in the training phase (when the system / algorithm is trained) and in the operational phase (when the system / algorithm is used for flash sound detection).

[0020] In another embodiment, one or more features related to the time, amplitude and frequency of the audio segments are used as characteristic features extracted from each audio segment.

[0021] A wide range of features are generally available as characteristic features of an audio segment, including one or more of the following: start time, end time, amplitude, gradient, instrumentation, keycode, melody, rhythm, pitch, energy, note offset, one or more Mel Frequency Cepstral Coefficients (MFCCs), Mel Spectrogram, Spectral Centroid, Spectral Flux, Zero Crossing Rate, Meter, Timbre, Pitch, Harmony, Root Mean Square Energy, Amplitude Envelope, Band Energy Ratio, Spectrogram, Constant Q Transform, Spectral Rolloff, Spectral Variance. In one implementation, selected MFCC features with 8-13 cepstral coefficients are used instead of the much larger number, e.g. 39 MFCC features, that are generally available for audio segments. The 0th coefficient, which represents the average logarithmic energy of the input signal carrying very little speaker specific information, is preferably not used, as it may lead to the detection of background noise (flash sounds are generally transient, lasting about 0.1 seconds). Thus, the present invention can effectively detect the low amplitude and short duration flash sound produced by photoepilation devices in the presence of various background noises, while providing the necessary low latency and memory.

[0022] The flash sound may be generated by triggering a light source such as a flash lamp, or may be artificially synthesized when a light source such as an LED is used. For example, during operation of a flash lamp or a xenon arc lamp, a sound (corresponding to a flash sound) is generated due to a pulse discharge. Laser radiation causes electrolysis in the air, which results in a sound. Similarly, sound is generated when an intense pulse of light comes into contact with the skin. However, other embodiments of the treatment device may use a semiconductor light source such as an LED, which does not generate a discharge sound due to a different mechanism. In that case, an additional sound generator may be used to synthesize a flash sound for each treatment light pulse.

[0023] According to another embodiment of the first aspect, the processing unit is configured to arrange the extracted features related to time, amplitude and frequency of the audio segments in an array and apply the array to a trained algorithm or computing system to detect, for each audio segment, whether the audio segment contains a flash sound. The format of the array can be configured, for example, such that the extracted features are arranged or sorted according to time-related features, amplitude-related features and frequency-related features, which makes it easier or even possible to process them with a trained algorithm or computing system, such as an XGBoost classification model.

[0024] Applying the trained algorithm or computing system to the audio segments by the processing unit includes constructing a classification boundary based on one or more characteristic features. XGBoost, for example, uses an ensemble tree-based technique to construct a classification boundary based on the input features passed to it. Parallel processing, tree pruning, and regularization are used here to avoid overfitting and provide a balanced model. Based on the classification boundary, each audio segment can be classified into either flash or non-flash so that the total number of flashes in the treatment can be predicted.

[0025] In another embodiment of the first aspect, the processing unit is configured to use the result of the detection for further training of an algorithm or computing system, such that during operation the detection result is not only used for detecting the flash sound but also at the same time for further training to further improve the trained algorithm or computing system.

[0026] According to a second aspect of the present invention, an audio signal sensed during a treatment procedure or artificially created is used for training an algorithm or a computing system. In the same way as is done during an actual treatment procedure, the audio signal is divided into audio segments and one or more characteristic features are extracted from the audio segments. From the annotation of the audio signal, flash sound information indicating whether the audio segment contains a flash sound is derived from the one or more characteristic features. In this context, the annotation of the audio signal shall be broadly understood as an indication of the flash sound in the audio signal, in particular the time information of the flash sound (e.g., timestamps or other indications indicating at what point in time the audio signal contains a flash sound, such as start time, end time, peak time, duration, etc.). The extracted characteristic features and the derived flash sound information are then used to train an algorithm or a computing system to detect flash sounds in the audio segment based on the one or more characteristic features.

[0027] The input unit is adapted to retrieve annotations of the audio signal, which annotations have been previously created by a user and have been stored, for example together with or separately from the audio signal. Alternatively or additionally, the annotations can be created automatically, for example by an algorithm or a computing system or another processing entity. For example, the processing unit may be configured to generate annotations of the audio signal, for example automatically or based on user input.

[0028] According to a third aspect of the present invention, a different invention is obtained, which is presented to a user. This information is displayed on various screens, where a first screen requests user input regarding the treatment operation, a second screen indicates the connectivity of the device or electronic device comprising the device with the treatment device, and a third screen indicates one or more of the following: the operating status of the skin treatment device, a flash sound counter that counts the number of flash sounds detected since the start of the treatment operation or a predetermined or user-instructed moment, an indication of whether the flash sound counter has exceeded a flash sound threshold, and an indication that the treatment operation has ended. This allows the user to know the current progress of the treatment operation and when to end the treatment so that over-treatment can be avoided or at least reduced.

[0029] The system according to the present invention comprises a treatment device, an audio sensor, and one or more of the above-mentioned devices. According to an embodiment of the system, the audio sensor, for example a microphone, is part of the device or electronic device comprising the device, in particular a computer, laptop, tablet, or smartphone (or a separate entity corresponding to the device only). In another embodiment, the device is part of the treatment device or electronic device, in particular a computer, laptop, tablet, or smartphone. In general, the above-mentioned aspects of the device may be implemented on the same device or on different devices. In a practical implementation, all aspects of the device and method are implemented on the same device, for example a tablet or smartphone.

[0030] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief description of the drawings]

[0031] [Figure 1] FIG. 1 illustrates an embodiment of an exemplary treatment device that can be used with the present invention. [Diagram 2] 1 is a schematic diagram of an exemplary embodiment of a system according to the present invention; [Diagram 3] 1 is a flow chart of an embodiment of a method according to a first aspect of the present invention. [Figure 4] 4 is a flow chart of an embodiment of a method according to the second aspect of the present invention. [Diagram 5] FIG. 2 is a visualization of an audio signal, with the flash sound and background noise indicated. [Figure 6] A zoomed-in visualization of two flash sounds. [Figure 7] FIG. 2 is a schematic diagram illustrating a pre-treatment according to the present invention. [Figure 8] FIG. 2 is a schematic diagram illustrating feature extraction according to the present invention. [Figure 9] 1 is a flow chart illustrating a model deployment to production process in accordance with the present invention. [Figure 10] 4 is a flow chart of an embodiment of a method according to the third aspect of the present invention. [Figure 11] 2A-2C are diagrams of various screens illustrating a user interface according to the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0032] FIG. 1 is a diagram of an exemplary treatment device 2 that can be used to apply light pulses (also called flashes) to an area of ​​skin. It will be understood that the treatment device 2 of FIG. 1 is provided merely as an example of a treatment device 2 that can be used with the present invention, and is not limited to the form shown in FIG. 1, or to being a handheld treatment device. The treatment device 2 is for use on the body of a subject (e.g., a human or animal) and is to be held in one or both hands of a user during use. The treatment device 2 is for performing some treatment treatment on the skin or body of a subject using one or more light pulses when the treatment device 2 is in contact with or near a body part of the subject. The treatment treatment can be the removal of unwanted body hair by laser and / or light therapy (known as photoepilation treatment or IPL (intense pulsed light) treatment). Alternatively, the treatment procedure may be laser and / or light therapy for reasons other than hair removal or growth prevention, such as dermatological (skin) treatment, treatment of acne, phototherapy treatment of other types of skin conditions, skin revitalization, skin tightening, or treatment of port wine stains.

[0033] As described herein, the treatment device 2 is operated or used by a "user," and the treatment device 2 is used on the body of a "subject." In some cases, the user and the subject are the same person, i.e., the treatment device 2 is held by the user in his / her hand and used on himself / herself (e.g., used on the skin of the user's leg). In other cases, the user and the subject are different people, e.g., the treatment device 2 is held by the user in his / her hand and used on someone else. In either case, there is little or no perceptible change to the skin upon or immediately after application of the light pulse, making it difficult for the user to completely cover the body site and / or to avoid over-treating certain areas of the body site.

[0034] The exemplary treatment device 2 comprises a housing 4 including at least a grip portion 5 and a head portion 6. The grip portion 5 is shaped to allow a user to grasp the treatment device 2 in one hand. The head portion 6 is at a head end 8 of the housing 4, and the head portion 6 is intended to be placed in contact with a subject to perform a treatment procedure on the subject's body or skin at a location where the head portion 6 is in contact with the body or skin.

[0035] The treatment device 2 is for performing a treatment procedure using light pulses. To this end, in FIG. 1, the head portion 6 comprises an aperture 10 disposed in or on the housing 4, such that the aperture 10 can be placed adjacent to or on (i.e. in contact with) the skin of a subject. The treatment device 2 includes one or more light sources 12 for generating light pulses to be applied to the skin of the subject via the aperture 10 to perform the treatment procedure. The one or more light sources 12 are disposed within the housing 4 such that the light pulses are provided from the one or more light sources 12 through the aperture 10. The aperture 10 is in the form of an opening in the head end 8 of the housing 4, or in the form of a window (including a waveguide) that is transparent or semi-transparent to the light pulses (i.e. the light pulses can pass through the window).

[0036] In the exemplary embodiment shown in Figure 1, the aperture 10 has a generally rectangular shape, resulting in a generally rectangular shaped skin treatment area on the skin. It will be understood that the aperture 10 may have any other desired shape. For example, the aperture 10 may be square, oval, circular, or any other polygonal shape.

[0037] 1, the treatment device 2 has one or more removable attachments for a head portion 6 of the treatment device 2, each of which includes an opening 10. The attachments are for use on different body areas and have different sized openings 10. For example, one attachment may be for use on larger body areas such as the legs and therefore have a larger opening 10, while another attachment may be for use on the skin / hair above the upper lip and therefore have a smaller opening 10.

[0038] The one or more light sources 12 may generate light pulses of any suitable or desired wavelength (or range of wavelengths) and / or intensity. For example, the light sources 12 may generate visible light, infrared (IR) light, and / or ultraviolet (UV) light. Each light source 12 may comprise any suitable type of light source, such as one or more light emitting diodes (LEDs), (xenon) flash lamps, one or more lasers, etc. The light sources 12 may provide light pulses having a spectral content in the 450-1200 nm range for a duration of approximately 2.5 milliseconds (ms), because these wavelengths are absorbed to heat melanin in the hair and hair roots, which causes the hair follicle to go dormant and inhibits hair regrowth.

[0039] The one or more light sources 12 are configured to provide a pulse of light. That is, the light source 12 is configured to generate high intensity light for a short duration (e.g., less than one second). The intensity of the light pulse should be high enough to perform a treatment procedure on the skin or body area adjacent the aperture 10.

[0040] The illustrated treatment device 2 optionally includes two skin contact sensors 14, 16 located on or in the head portion 6, which are used to determine whether the head portion 6 is in contact with the skin. The skin contact sensors 14, 16 measure a parameter indicative of whether the head portion 6 is in contact with the skin and generate a corresponding measurement signal including a time series of measurements of the parameter. The measurement signal can be processed to determine whether the head portion 6 is in contact with the skin. Typically, skin contact sensors are used in treatment devices 2, particularly photoepilators, to ensure that the treatment device 2 is in proper contact with the skin before a light pulse is generated, in order to avoid the light pulse being directed into the user's or subject's eyes.

[0041] In some embodiments, the parameter may be capacitance, in which case the skin contact sensors 14, 16 may measure capacitance via corresponding pairs of electrical contacts or electrodes on the surface of the head portion 6, with the measured capacitance indicating whether skin contact is present. In alternative embodiments, the parameter may be light intensity or level, in which case the skin contact sensors 14, 16 may be light sensors that measure the intensity or level of light incident on the light sensors, with the measured intensity or level indicating whether skin contact is present (e.g., less / no light may indicate skin contact, as the skin obscures the light sensors 14, 16, and vice versa). In other alternative embodiments, the parameter may be a measure of contact pressure, in which case the skin contact sensors 14, 16 may measure contact pressure via corresponding pressure sensors or mechanical switches, with the measured contact pressure indicating whether skin contact is present.

[0042] The illustrated treatment device 2 includes an optional skin tone sensor 18 located on or in the head portion 6 that can be used to determine the skin tone of the skin with which the head portion 6 is in contact. The skin tone sensor 18 can measure a parameter indicative of the skin tone of the skin and generate a measurement signal that includes a measurement of a time series of the parameter. The measurement signal can be processed to determine the skin tone of the skin with which the head portion 6 is in contact. Typically, a skin tone sensor is used in a treatment device 2, particularly a photoepilator, to ensure that the light pulses have an appropriate intensity for the type of skin being treated, or even to prevent light pulses from being generated when the skin type is not suitable for light pulses (e.g. darker skin with a much higher melanin content).

[0043] In some embodiments, the optional skin tone sensor 18 may be an optical sensor, and the parameter measured by the optical sensor may be the intensity or level of light at a particular wavelength or wavelengths reflected from the skin. The measured intensity or level of reflected light at a particular wavelength may be indicative of skin tone. The measured intensity or level of reflected light may be based on the concentration of melanin in the skin, and thus the measured intensity or level may be indicative of melanin concentration. Melanin concentration may be derived, for example, from measurements of light reflectance at 660 nm (red) and 880 nm (infrared) wavelengths.

[0044] The illustrated treatment device 2 also includes user controls 20 that are operable by a user to operate the treatment device 2 such that the head portion 6 performs a desired treatment action on the subject's body (e.g., generation of one or more light pulses by the one or more light sources 12). The user controls 20 may be in the form of switches, buttons, touch pads, etc.

[0045] As noted above, one or more sensors are optionally used to provide information regarding the position (including orientation) and / or movement of the treatment device over time so that appropriate feedback can be determined, and in some embodiments, the one or more sensors can be part of or within the treatment device 2. If the sensor is one or more motion sensors, such as an accelerometer, gyroscope, etc., the one or more motion sensors can be part of, e.g., internal to, the treatment device 2 (and therefore not shown in FIG. 1). In embodiments in which the one or more sensors include an imaging unit (e.g., a camera), the imaging unit can be on or within the treatment device 2. As shown in FIG. 1, the imaging unit 22 is positioned on the housing 4 of the treatment device 2 so as to capture images or video sequences of the skin or body area in front of the treatment device 2. That is, the imaging unit 22 is positioned on the housing 4 of the treatment device 2 so as to capture images or video sequences from a direction generally parallel to the direction of light pulses emitted from the light source 12 through the opening 10. In some implementations, the imaging unit 22 is positioned on the housing 4 such that the imaging unit 22 cannot see the opening 10 and the portion of the skin in contact with the head end 6 when the head portion 6 is brought into contact with the subject so that a light pulse can be applied to the portion of the skin visible from the light source 12 through the opening 10. In embodiments in which the imaging unit 22 is positioned on or within the housing 4 of the treatment device 2, the imaging unit 22 is in a fixed configuration relative to the rest of the treatment device 2.

[0046] The treatment device 2 further comprises an audio source 24 (sound generator), such as a speaker, for emitting a short audio pulse (also called a flash sound) each time a light pulse (flash) is generated. This is particularly true when the light source 12 is a semiconductor source, such as an LED or VCSEL, that does not generate the characteristic arc lamp pulse discharge flash sound. In other embodiments, the flash sound may be generated (e.g. automatically) by the light source when the light pulse is generated. For example, a sound pulse, such as a tick sound, may be generated each time a light pulse is emitted. The flash sound is, for example, a short audio pulse of a particular frequency (e.g. 9 kHz or any other selected frequency) or frequency range (e.g. within the range of 1-10 kHz) and of a short duration (e.g. within the range of 0.1-2 seconds, e.g. about 1 second) that allows the user to hear the flash sound.

[0047] Fig. 2 is a schematic diagram of an exemplary system 40 according to the present invention, comprising a treatment device 2 and an apparatus 42 for use with the treatment device 2, for controlling the treatment device 2, e.g. as shown in Fig. 1. The system 40 further comprises an audio sensor 44 (e.g. a microphone, but there may be multiple audio sensors) configured to sense audio signals during the treatment procedure, which is preferably part of the apparatus 42, but may also be an external element. Furthermore, the system 40 optionally comprises a feedback unit 46 (e.g. a graphical user interface or an audio output) for providing feedback to the user and / or the subject, especially regarding the progress of the treatment procedure or any instructions to be followed by the user.

[0048] As noted above, in some embodiments, the device 42 may be a separate device from the treatment device 2, in which case the device 42 may be in the form of an electronic device such as a smart phone, a smart watch, a tablet, a personal digital assistant (PDA), a laptop, a desktop computer, a remote server, a smart mirror, etc. In other embodiments, the device 42, and in particular the functionality of the present invention provided by the device 42, is part of the treatment device 2.

[0049] The device 42 includes a processing unit 50 that generally controls the operation of the device 42 and enables the device 42 to perform the methods and techniques described herein. The processing unit 50 is further configured to receive measurement signals from the audio sensor 44, process the measurement signals to determine feedback to be provided for the treatment device, and provide feedback control signals to the feedback unit 46. Thus, the processing unit 50 is configured to receive measurement signals from the audio sensor 44 directly, in embodiments where the sensor 44 is part of the device 42, or via another component, in embodiments where the sensor 44 is separate from the device 42. In either case, the processing unit 50 may include or comprise one or more input ports or wirings for receiving the measurement signals via the input unit 48. The processing unit 50 may also include or comprise one or more output ports or wirings for outputting the feedback control signals via the output unit 52.

[0050] The processing unit 50 can be implemented in many ways using software and / or hardware to perform the various functions described herein. The processing unit 50 comprises one or more microprocessors or digital signal processors (DSPs) that are programmed using software or computer program code to perform the required functions and / or control the components of the processing unit 50 to achieve the required functions. The processing unit 50 is implemented as a combination of dedicated hardware (e.g., amplifiers, preamplifiers, analog-to-digital converters (ADCs), and / or digital-to-analog converters (DACs)) to perform some functions and processors (e.g., one or more programmed microprocessors, controllers, DSPs, and associated circuitry) to perform other functions. Examples of components employed in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, DSPs, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), hardware for implementing neural networks and / or so-called artificial intelligence (AI) hardware accelerators (i.e., processors or other hardware designed specifically for AI applications that can be used in parallel with a main processor), and / or hardware.

[0051] The processing unit 50 optionally comprises or is associated with a memory unit (not shown). The memory unit may store data, information, and / or signals (including images) used by the processing unit 50 in controlling the operation of the device 42 and / or in implementing or executing the methods described herein. In some implementations, the memory unit stores computer-readable code executable by the processing unit 50 to cause the processing unit 50 to perform one or more functions, including the methods described herein. In certain embodiments, the program code may be in the form of an application for a smartphone, tablet, laptop, computer, or server. The memory unit may comprise any type of non-transitory machine-readable medium, such as a cache or system memory, including volatile and non-volatile computer memory, such as random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), and electrically erasable PROM (EEPROM), and the memory unit may be implemented in the form of a memory chip, an optical disk (e.g., a compact disk (CD), a digital versatile disk (DVD), or a Blu-ray disk), a hard disk, a tape storage solution, or a solid-state device, including a memory stick, a solid-state drive (SSD), a memory card, etc.

[0052] In the embodiment shown in Fig. 2, the device 42 is shown as being separate from the sensor 44, and therefore includes an input unit 48, such as an interface circuit, to enable the device 42 to receive measurement signals from the sensor 44. The interface circuit 48 in the device 42 allows data connection and / or data exchange with other devices, including any one or more of the sensor 44, the treatment device 2, a server, and a database. The connection to the sensor 44 (or any electronic device, such as the treatment device 2) may be direct or indirect (e.g., via the Internet), in which case the interface circuit 48 may enable a connection between the device 42 and a network, or a direct connection between the device 42 and another device (e.g., the sensor 44 and / or the treatment device 2) via any desired wired or wireless communication protocol. For example, the interface circuit 48 may operate using WiFi, Bluetooth, Zigbee, or any cellular communication protocol, including but not limited to Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Long Term Evolution (LTE), LTE-Advanced, etc. In the case of a wireless connection, the interface circuit 48 (and thus the device 42) includes one or more antennas suitable for transmission / reception over a transmission medium (e.g., air). Alternatively, in the case of a wireless connection, the interface circuit 48 includes means (e.g., a connector or plug) for allowing the interface circuit 48 to be connected to one or more suitable antennas external to the device 42 suitable for transmission / reception over a transmission medium (e.g., air). The interface circuit 50 is connected to the processing unit 50.

[0053] Furthermore, the device 42 comprises an output unit 52, such as an output circuit, for enabling the device to output information, inter alia, to the feedback unit 46. The output unit may generally be configured to have similar or the same functionality as the input unit 48.

[0054] The feedback unit 46 includes one or more components that enable a user of the device 42 to input information, data, and / or commands to the device 42 (e.g., via the input unit 48) and / or enable the device 42 to output information or data to a user of the device 42 (via the output unit 52). The user interface may comprise any suitable input components, including, but not limited to, a keyboard, a keypad, one or more buttons, switches, or dials, a mouse, a trackpad, a touch screen, a stylus, a camera, a microphone, etc., and the user interface may comprise any suitable output components, including, but not limited to, a display unit or display screen, one or more lights or lighting elements, one or more speakers, a vibration element, etc.

[0055] In some embodiments, the feedback unit 46 is attached to or is part of the treatment device 2. In other embodiments, the feedback unit 46 is part of the apparatus 42. In these embodiments, the feedback unit 46 may be part of or utilize any of the output components of the user interface components of the apparatus 42. In further embodiments, the feedback unit 46 may be separate from the treatment device 2 and the apparatus 42.

[0056] It will be appreciated that an actual implementation of device 42 may include components in addition to those shown in Figure 2. For example, device 42 may also include a power source, such as a battery, or components to enable device 42 to be connected to a mains power source.

[0057] The present invention has various aspects, in particular the actual use of the apparatus 42 during an actual treatment procedure and the training of the apparatus 42 to reliably recognize the flashing sound emitted by the treatment device 2 during a treatment procedure.

[0058] A flow chart of an embodiment of the method 100 according to the first aspect, used during the actual operation of the treatment device, is shown in Fig. 3. In a first step 101 of the method 100, an audio signal sensed (by the audio sensor 44) during the treatment procedure is acquired (by the input unit 48), i.e. received or retrieved from the audio sensor 44. In a second step 102, the acquired audio signal is segmented into audio segments, in particular into audio segments of a given or pre-determined duration (e.g. 1 second or any other suitable duration). In another (combinable) embodiment, said audio segments of said pre-determined duration are acquired directly by the input unit 48 during the treatment procedure. In a third step 103, for each audio segment, one or more characteristic features (described in more detail below) are extracted from the audio segment. Then, in a fourth step 104, for each audio segment, it is detected whether the audio segment contains a flash sound. This is done by applying a trained algorithm or computing system, which has been trained on a plurality of audio segments each including one or more characteristic features, to detect flash sounds in the audio segments. In a fifth step 105, each time a flash sound is detected, a flash sound counter is incremented, which counts the number of flash sounds detected since the start of the treatment procedure or a predetermined or user-instructed moment. Finally, in step 106, information is output to the feedback unit 46 via the output unit 52. The information includes one or more of the flash sound counter (i.e. its current value), an indication of whether the flash sound counter has exceeded a flash sound threshold, and an indication that the treatment procedure has ended.

[0059] The flash sound threshold may, for example, be preset and may reflect the maximum number of flashes required or appropriate or prescribed for a particular body area to be treated. Such settings may be stored in the device 42 or may be entered by the user, who may take the settings from a table, for example in a user manual, or may select or calculate the settings manually or with the support of a feedback unit or user guidance on the device 42 or treatment device 2. For example, the flash sound threshold may depend not only on the size of the treatment area, but also on the skin type, skin color, number of previous treatments, time since the last treatment, etc.

[0060] A flow chart of an embodiment of the method 200 according to the second aspect, used for training, is shown in FIG. 4. In a first step 201 of the method 200, an audio signal is acquired, sensed (by the audio sensor 44) during a treatment procedure or artificially created (e.g. by the sound processor of the device or an external device). The audio signal is in this case a "real" audio signal as recorded during a treatment session or a simulated one, and thus includes multiple flash sounds (and possibly background noise). The length of the audio signal may thus be from a few seconds to tens of seconds or even minutes. Furthermore, the audio signal includes annotations, which may be generated by a user while listening to the audio signal, or may be generated automatically by an algorithm or device that artificially created the audio signal (and generated the annotations at the same time). The annotations of the audio signal include flash sound information indicating the flash sounds in the audio signal, in particular time information of the flash sounds, such as the start, end, highest peak, duration, etc. of the flash sounds.

[0061] In a second step 202, the (annotated) audio signal is segmented into audio segments, in particular into audio segments of a predefined or pre-determined duration. In a third step 203, for each audio segment, one or more characteristic features, e.g. a set of MFCC coefficients (described in more detail below), are extracted from the audio segment. In a fourth step 204, from the annotations of the audio signal, flash sound information is derived for each audio segment indicating whether the audio segment contains a flash sound, e.g. the flash sound information is read from a file or metadata representing an annotation associated with or embedded in the (annotated) audio signal. In a fifth step 205, the extracted one or more characteristic features and the derived flash sound information are provided to an algorithm or computing system for training the algorithm or computing system to detect flash sounds in the audio segment based on the one or more characteristic features, for each audio segment.

[0062] There are several machine learning and deep learning architectures for audio classification in the public domain. One such use case is the public dataset UrbanSound8K. This dataset contains 8732 labeled sounds extracted from 10 classes of urban sounds: car horns, dogs barking, gunshots, children playing, air conditioners, excavation, car idling, jackhammers, sirens, and street music. Mel-frequency cepstral coefficients (MFCCs) are used for feature extraction, which is then zero-padded and the input is passed to a convolutional neural network (CNN) model. Using MFCC features, a training accuracy of 98.12% and a validation accuracy of 91.23% are achieved. To improve the accuracy, feature padding is applied instead of zero padding, and other features such as zero-crossing rate, pitch, RMSE, and chroma are used along with MFCCs. After 100 epochs of training, a training accuracy of 99.21% and a validation accuracy of 97.29% are achieved.

[0063] Another dataset is ESC-50, which consists of 50 classes. This dataset consists of 2000 audio pieces, each with a duration of 5 seconds. The performance on ESC-50 is 94% accurate. We use the audio spectrogram transformer and utilize a convolution-free architecture for audio classification. Using the audio spectrogram transformer, we achieve a training accuracy of 94.7%.

[0064] TUT Rare Sound Events 2017 consists of sound events personalized for each target class, and recordings of everyday acoustic scenes that serve as backgrounds. The target sound event categories are babies crying, glass breaking, and gunshots. We organized several events to solve the above challenges, and achieved the highest accuracy of -93% on the evaluation dataset and -96% on the development dataset. The proposed method to solve this problem is divided into four steps: feature extraction (log magnitude and mel spectrogram), transformation of spectral features with ID convNet, time dependency with RNN-LSTM, and post-processing to detect the presence and occurrence time of audio events.

[0065] Uyen Diep (Sally), "Classification of Cough Events", ECE Senior Capstone Project 2017, Tech Notes, presents a cough analysis as part of an automatic patient cough monitoring system, which is divided into two main stages: event detection and event classification. In this method, sound events are divided into overlapping time frames of 32ms or 16ms length. Features are calculated for each frame and fed into a classifier to determine whether the frame has characteristics of a cough. The results are then combined to make an overall decision about the event. MFCCs are used in this method.

[0066] Known methods for audio classification based on signal processing are mainly multi-class classification using deep neural networks. The sounds that are mainly classified last for several seconds and are continuous in nature. In contrast, in the flash detection in the present scenario of photoepilation, the flash generally lasts for several milliseconds and is non-periodic. In one embodiment, a classical machine learning method is used to detect and train the flash sound to build a binary classification model to detect whether a flash is present or not. The significance of using such a machine learning (ML) method rather than deep learning is that the size of the model is optimized, which makes it easier to deploy the model in an app.

[0067] For example, when performing hair removal treatment at home, there are several challenges for high quality hair removal treatment, proper recommendation, and frictionless experience. In particular, the sound pattern of the treatment device needs to be detected. The sound pattern signal consists of the flash sound emitted during the use of the treatment device, but there may be potential background noises such as, for example, TV sounds, radio, baby crying, traffic noise, people talking, doors opening and closing, etc. Therefore, it is necessary to accurately identify the intended sound pattern in the audio signal. The present invention utilizes signal processing and a trained algorithm or computing system, in particular machine learning (ML), with a suitable set of features (e.g., 12-13 selected MFCC features) to reliably and accurately detect the flash sound even in the presence of some background noise. The trained algorithm or computing system, for example, an ML model (such as XGBoost), learns to distinguish between background noise (non-flash) and flash sound during the training process. In this case, the annotation helps the model understand the patterns that comprise the flash and non-flash audio segments.

[0068] The present invention therefore addresses several problems, including one or more of the following: Ability to accurately detect sounds of various amplitudes from audio captured at distances from 50cm to 1m -Immune to various background noises such as television, opening and closing doors, running water, radio, etc. Automatic detection of the exact number of flashes (used during a treatment) Techniques to reduce annotation effort Separation of various background noises (television, dripping water, etc.) that overlap with the flash sound Accurate detection to process the signal and detect only the flash sound The ability to have a lightweight, low-latency architecture for embedding Artificial Intelligence (AI) algorithms into both traditional applications and operating systems.

[0069] Unlike prior art approaches such as the cough detection method disclosed by Uyen Diep in the above mentioned publication, the present invention allows reliable detection of (artificial) flash sounds emitted by a treatment device in the presence of background noise, since the characteristics of the flash sounds are significantly different from the characteristics of a human cough. The problem solved by the present invention can be seen as detecting the exact number of flash sounds occurring during a treatment. Flash sounds are not continuous in nature, lasting for a fraction of a second, i.e., they are aperiodic and transient (as opposed to aperiodic continuous sounds and periodic (simple or complex) sounds). The raw audio signals acquired during treatment of different body areas of a subject, such as legs and armpits, are of different duration. For example, the length of the armpit audio varies from 30 seconds to 105 seconds, and the length of the leg audio varies from 30 seconds to 180 seconds. The audio signal includes flash sounds with a minimum interval of, for example, 0.8 seconds. The audio signal also includes background noise.

[0070] As a result, the processing technique of the present invention differs from the prior art. For example, sound events such as coughs can be grouped into overlapping frames, but in order to count the exact number of flashes, the technique splits the audio signal into non-overlapping audio segments during pre-processing. Non-overlapping audio segments are particularly useful to avoid double counting of flashes. MFCCs typically use only 8-13 cepstral coefficients. The 0th coefficient, commonly used in the prior art, for example disclosed by Uyen Diep, represents the average logarithmic energy of the input signal and carries very little speaker-specific information. In the preferred embodiment of the present invention, the 0th coefficient is not used, since it leads to the detection of background noise. The present invention thus provides the ability to detect low amplitude, short duration pulses produced by treatment devices with various background noises. MFCC is simply a feature extraction step and the right combination of these features with an appropriate algorithm / computational system (e.g., an appropriate ML model) contributes to achieving the desired effect and solving the addressed problem, in particular achieving low latency, low memory footprint, and high accuracy due to the ability to distinguish subtle flash sounds from various background noises with proper feature selection.

[0071] In an exemplary implementation, data (audio signals) are collected by a microphone on a smartphone located close to the treatment device (e.g., within 0.5-2.5 m). The audio signals include the flash sound and the inherent background noise.

[0072] During training, whenever a flash sound is present, the raw audio signal can be annotated, i.e., an indication is added, such as a comment or metadata, indicating the moment in time in the audio signal where the flash sound is present. Pre-processing of the audio signal is performed by quantizing the raw audio signal or the annotated signal into segments of a given duration, e.g., one second chunks, and labeling the segments as ticks (i.e., flash sounds) and non-ticks (non-flash sounds). In addition, background noise filtering can be performed and time-domain and frequency-domain features, such as MFCCs, mel spectrograms, can be extracted from the audio signal to create an audio feature array. In one embodiment, the feature array is used to build an ML model, which has the ability to detect flash sounds in near real-time. The ML model can be deployed, for example, as a program file into an app (application or software program), which allows the model to be run on a desired device, such as a smartphone.

[0073] While generally any algorithm or computing system can be used, and generally any kind of learning system, machine learning, classification model, binary classification model, neural network, or convolutional neural network can be used as the trained algorithm or computing system, in one embodiment, an XGBoost binary classification model is constructed to detect whether a flash sound is present in an audio signal. Techniques for constructing such an XGBoost classifier are generally known in the art. An exemplary XGBoost classifier may be constructed with the following hyper-parameters: leaming_rate=0.02, n_estimators=600, objective='binary: logistic'. Other parameters and other classifiers or ML models may also be used. It is noted that in this context, generally, a specific implementation or architecture or layout of the trained algorithm or computing system is not required, and for example, a standard learning system or neural network with a known or conventional architecture may be used for the task of detecting a flash sound.

[0074] Annotations of the audio signal (as shown in FIG. 4) are used for training. An embodiment of a technique for creating such annotations is described below. In the implementation described above, audio data is collected by a smartphone microphone, e.g. in real time or with temporary storage, preferably at a distance of 0.5 m to 2.5 m from the photoepilation device. Flash sounds are extracted by the ML model. The extracted data includes the start and end time of each flash sound in the audio file. In general, any annotation tool that allows adding annotations to audio files can be used for this purpose. In an exemplary implementation, the Applicant's annotation tool "Barista" is used.

[0075] In one embodiment, a raw audio file is opened, the peaks are visualized, and the audio is listened to (e.g., once). A visualization of a portion of such an audio file is shown in FIG. 5, where a flash sound ("flash tick") and a background noise ("door opening or any other noise") are indicated. The audio file is then started to play, and the playback is paused where the flash sound can be visualized or heard. The audio file is then zoomed in on that area. FIG. 6 shows such a zoomed in visualization of one second of two flash sounds, with the start and end times indicated. An area to be annotated can then be selected. The selected area is played back and listened to for cross-validation. This process is repeated until the end of the audio file.

[0076] In an embodiment, as illustrated generally in FIG. 7, audio pre-processing 300 of a raw audio file 301 and / or annotated audio file 302 (with indication of flash sounds) is performed by quantizing the audio file into audio segments (chunks) 303 of a certain length (e.g. in the range of 0.1-5 seconds, e.g. 0.5 seconds or 1 second). Each audio segment is then labeled as a flash sound or a non-flash sound. Furthermore, basic audio operations 304 can be performed, such as repeating the pre-processing multiple times, changing the pitch of the audio signal of the audio file, etc. This is done because raw audio signals are not always of sufficient or optimal quality, and may be corrupted, for example, by background noise.

[0077] In an embodiment, feature extraction 400 is performed to analyze audio data, as shown diagrammatically in FIG. 8. Audio data can generally be considered as a 3D signal representing time, amplitude, and frequency. Feature extraction techniques are utilized because the 3D format cannot (or is not easy to) be directly processed by an AI model. Before passing the audio signal to the XGBoost classification model, the signal is converted into an array format, so that the signal is ready for further processing and analysis by the AI ​​model. Generally, in a first step 401, pre-emphasis, frame blocking, and windowing are performed, and then in a second step 402, an FFT is applied. Then, a filter bank 403, e.g. a mel-scale filter bank, a logarithm 404, and again a filter bank 405 are applied to obtain the desired features, e.g. MFCC features 406.

[0078] MFCC refers to the Mel scale, a scale associated with the perception of pitch. From the MFCC features, several coefficient values ​​are obtained that describe the characteristics of an audio segment, i.e., a sound. In one embodiment, in a first step 401, a given audio signal is divided into short frames (audio segments). This step is performed because the frequency of the signal changes over time, and therefore the frequency contour of the audio signal is lost over time when computing the Fourier transform of the entire audio signal. Furthermore, a windowing process (e.g. with a Hamming window) is performed to cross-check the data for infinity and to reduce spectral leakage.

[0079] In a second step 402, a discrete Fourier transform is calculated to obtain a simplified version of the Fourier transform and to obtain real-valued coefficients (e.g., a total of 39 coefficients per frame). In a third step 403, a Mel-spaced filter bank is applied, which includes converting the lowest and highest frequencies to Mel, creating equally spaced points as several Mel bands to use, converting the points back to Hertz, rounding the points to the nearest frequency bin (since there is a discrete signal), and creating a triangular filter. As a result, 39 coefficients of the MFCC features are obtained.

[0080] In general, every audio signal is composed of multiple features. However, it is necessary to extract the characteristics that are relevant to the problem to be solved. These audio features can be used to build and train an intelligent audio classification model.

[0081] In an embodiment of the present invention, Mel-frequency Cepstral Coefficients (MFCCs) are used as features. The MFCCs of a signal are a small set of features (usually around 10-20) that succinctly describe the overall shape of the spectral envelope. MFCCs model the characteristics of the human voice. Typically, there are around 39 coefficients per frame. In one embodiment, only the first coefficients are used, e.g., the first 12-13 coefficients (preferably the first 13 coefficients) not including the 0th component, which represents the average logarithmic energy of the input signal, because the first few coefficients encode most of the information (e.g., spectral envelope, formants, etc.). MFCC features perform well in music and speech processing, but lack robustness to noise.

[0082] There are several audio categorizations based on the categories indicated in the table below.

[0083] [Table 1]

[0084] A categorization of audio features can be done based on the following areas:

[0085] [Table 2]

[0086] After creating audio segments in a pre-processing step, the extracted MFCC features are passed to the model.

[0087] In the following, an embodiment of the AI ​​model training is described in more detail. The input to the model can be a matrix (or vector) with, for example, 13 columns as MFCC coefficients and a number of rows corresponding to the length of the audio file (e.g., one row for each audio segment of, for example, 1 second in length).

[0088] In one embodiment, the following 13 MFCC coefficients are used for each audio segment of length 1 second (i.e., these 13 MFCC coefficients represent the 13 columns of a matrix): mfcc_second_derivative_std_1 mfcc_second_derivative_std_2 mfcc_second_derivative_std_3 mfcc_second_derivative_std_4 mfcc_second_derivative_std_5 mfcc_second_derivative_std_6 mfcc_second_derivative_std_7 mfcc_second_derivative_std_8 mfcc_second_derivative_std_9 mfcc_second_derivative_std_10 mfcc_second_derivative_std_11 mfcc_second_derivative_std_12 mfcc_second_derivative_std_13.

[0089] Note that these are common names commonly known to those skilled in the art. These features are selected and ranked based on their predictive power for identifying flash sounds. Here, the coefficient "mfcc_second_derivative_std_1" is the standard deviation of the second derivative of the MFCCs (also called the delta-delta coefficient). To arrive at the second derivative, MFCC features are extracted and then the second derivative on those MFCC features is computed.

[0090] Delta MFCCs are computed for each frame. For each frame, this is the current MFCC value minus the previous MFCC frame value. Since the differencing operation is sensitive to noise, this is similar to applying a smoothing filter, leading to improved accuracy. These matrix coefficients are then fed into the model along with the target features (during training).

[0091] Exemplary values ​​for the 13 coefficients for an exemplary audio segment of length 1 second are as follows: 2.17, 2.14, 1.34, 1.60, 1.94, 1.71, 1.62, 1.67, 1.66, 1.24, 1.48, 1.66, 1.65.

[0092] An example matrix might look like this (the first 13 columns are the feature vector (X), the last column (ticks) is the target variable (Y), and each row represents the values ​​of a different audio segment):

[0093] [Table 3]

[0094] An exemplary model used is the XGBoost classifier for classifying whether a flash sound is present or not in each audio segment of length, for example, 1 second, based on the extracted MFCC features. In general, ensemble models perform better in classification in regression tasks than non-ensemble models, such as linear regression, logistic regression, etc. XGBoost is an ensemble machine learning algorithm that uses a decision tree and gradient boosting framework. It can be used to solve multiple regression problems, classification problems, ranking problems, and user-defined prediction problems. Thus, the output of the MFCC operation, a matrix (rows with values ​​for audio segments), is fed to the ML model for classification. During training, each frame is annotated as being flash or not flash and passed as a target variable along with the feature matrix. Although generally different coefficients (i.e., features) are used in training and in actual use, it is preferable to use the same coefficients in training and in actual use of the device and method.

[0095] The audio files contain the treatment sounds, including a flash sound mask with device fan noise, along with other background noises. We collect audio files of many treatments and use them for training the AI ​​model and others for validation purposes. The duration of the armpit audio varies from 30 seconds to 105 seconds, and the duration of the leg audio varies from 30 seconds to 180 seconds. Using the above algorithm, a training accuracy of 96% and a validation accuracy of 90% are achieved.

[0096] The flash detection model is deployed to various operating systems, e.g., iOS, Android, etc. Figure 9 shows a flowchart of the process from model deployment to production. Retraining the model with new audio, taking into account the base model, is included as an important step. It is foreseen to fine-tune the model and its effectiveness using new datasets and save the best AI model. Model updating is an iterative process that includes continuous or iterative retraining, fine-tuning the best model, and deploying it to the mobile application.

[0097] During production, the microphone on a smartphone (or other useful device) is used to capture real-time audio as input and pass it to the latest model for flash detection.

[0098] A flow chart of an embodiment of a method 500 according to a third aspect used for providing a user interface is shown in Fig. 10. The method 500 is generally configured for controlling a user interface in a first step 501 to display a first screen requesting user input regarding a treatment procedure. In a second step 502, a user input is obtained via the first screen, said user input comprising an image of the subject's body part to be treated and / or size information of the size of the subject's body part to be treated. In a third step 503, a second screen is displayed indicating connectivity between the apparatus or an electronic device comprising the apparatus and the treatment device. In a fourth step 504, a third screen is displayed indicating at least one of the following: the operating state of the skin treatment device; a flash sound counter for counting the number of flash sounds detected since the start of a treatment procedure or a predetermined or user-initiated moment; an indication of whether the flash sound counter has exceeded the flash sound threshold; and An indication that a treatment procedure has been completed.

[0099] The corresponding user interface is shown at various stages in Fig. 11. Screen 601 (representing the first screen in step 501) guides the user on how to effectively receive the treatment and asks the user for the size of the desired treatment area, e.g. leg length. Screen 602 shows the user having entered the leg length (in step 502) and some number of flashes that would be ideal for a proper treatment will be displayed. Clicking continue takes the user to screen 603 where he is asked to click the start button to begin the treatment. Once the treatment has begun, screen 604 (representing the second screen in step 503) asks for access to (or, more generally, to indicate a connection to) the microphone. If this is granted, the detection of the flash sound will begin as shown in screen 605 and the total number of flashes used throughout the treatment will be incremented as shown in screen 606 (representing the third screen in step 504). In another embodiment, the number of flashes remaining required to treat the desired treatment area may be displayed, or both the total number of flashes applied and the number of flashes remaining may be displayed.

[0100] The screens shown in Figure 11 should be understood as examples: other information may be displayed on each screen, and other designs and numbers of screens may be present.

[0101] The present invention provides the advantage of automating the flash count for each treatment, which will help the user to count for subsequent treatments. Usually, one treatment requires about 30-40 flashes. Currently, the user has to manually track the flash count, which leads to a less satisfying user experience. In this case, a proper recommendation for the device will remove the friction in using the device (i.e., the flash count is automated without the user having to count manually during the treatment), resulting in a better user experience and satisfaction. The presented method was tested with various background noises and a model with the ability to identify the flash sound. Whenever a treatment does not meet the expected level, the treatment can be checked to see if the guidance or recommendation was followed.

[0102] While the present invention has been illustrated and described in detail in the drawings and the foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive, and the invention is not limited to the disclosed embodiments. In practicing the invention as claimed, those skilled in the art can understand and effect other variations of the disclosed embodiments, from a study of the drawings, the disclosure, and the appended claims.

[0103] In the claims, the words "comprise, include, have" do not exclude other elements or steps, and the singular elements do not exclude a plurality. A single element or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

[0104] The computer program may be stored / distributed on a suitable non-transitory medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.

[0105] Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. 1. An apparatus for use with a treatment device for performing a treatment action on a body part of a subject, the treatment device being configured to apply light pulses to skin of the body part to perform the treatment action, and to generate a flashing sound for each light pulse generated by the treatment device, the apparatus comprising: an input unit for acquiring an audio signal sensed during the treatment operation; a processing unit and an output unit, The processing unit Segmenting the audio signal into audio segments, in particular audio segments of predetermined or pre-determined duration; extracting one or more characteristic features for each audio segment from the audio segments; detect, for each audio segment, whether the audio segment includes a flash sound generated by the treatment device for each light pulse by applying a trained algorithm or computing system that has been trained with a plurality of audio segments, each audio segment including one or more of the characteristic features, to detect within the audio segment a flash sound generated by the treatment device for each light pulse; Varying a flash sound counter that counts the number of flash sounds detected since the start of the treatment operation or a predetermined or user-instructed moment; The output unit said flash sound counter; an indication of whether the flash sound counter has exceeded a flash sound threshold; and Indication that the treatment operation has ended 4. An apparatus for outputting one or more of:

2. the processing unit uses a learning system, machine learning, classification model, binary classification model, neural network, or convolutional neural network as a trained algorithm or computing system; 10. The apparatus of claim 1.

3. the processing unit extracts one or more features related to time, amplitude, and frequency of the audio segment as characteristic features; 3. The device according to claim 1 or 2.

4. the processing unit extracts one or more of the following characteristic features of the audio segment: start time, end time, amplitude, gradient, instrumentation, keycode, melody, rhythm, pitch, energy, note offset, one or more Mel Frequency Cepstral Coefficients (MFCC), Mel Spectrogram, spectral centroid, spectral flux, zero crossing rate, meter, timbre, pitch, harmony, root mean square energy, amplitude envelope, band energy ratio, spectrogram, constant Q transform, spectral roll off, spectral variance; 4. The apparatus of claim 3.

5. the processing unit arranges the extracted features related to time, amplitude, and frequency of the audio segments into an array and applies the array to the trained algorithm or computing system to detect, for each audio segment, whether the audio segment contains a flash sound; 5. The device according to claim 3 or 4.

6. applying the trained algorithm or computing system by the processing unit to the audio segment includes constructing a classification boundary based on the one or more characteristic features.

6. An apparatus according to any one of claims 1 to 5.

7. the processing unit uses the results of the detection for further training of the trained algorithm or computing system.

7. An apparatus according to any one of claims 1 to 6.

8. 1. An apparatus for use with a treatment device for performing a treatment action on a body part of a subject, the treatment device being configured to apply light pulses to skin of the body part to perform the treatment action, and to generate a flashing sound for each light pulse generated by the treatment device, the apparatus comprising: an input for receiving an audio signal sensed or artificially generated during said treatment operation; a processing unit; The processing unit Segmenting the audio signal into audio segments, in particular audio segments of predetermined or pre-determined duration; extracting one or more characteristic features for each audio segment from the audio segments; deriving, for each audio segment, flash sound information from annotations of the audio signal indicating whether the audio segment includes a flash sound generated by the treatment device for each light pulse, the annotations of the audio signal indicating within the audio signal the flash sound generated by the treatment device for each light pulse, in particular time information of the flash sound generated by the treatment device for each light pulse; The apparatus provides the extracted one or more characteristic features and the derived flash sound information to an algorithm or computing system to train the algorithm or computing system to detect, for each audio segment, a flash sound generated by the treatment device for each light pulse within the audio segment based on the one or more characteristic features.

9. 9. The device according to claim 8, wherein the input unit obtains the annotation of the audio signal and / or the processing unit generates the annotation of the audio signal, particularly automatically or based on user input.

10. displaying a first screen requesting user input regarding the treatment operation; acquiring the user input via the first screen, the user input including an image of the body part of the subject receiving treatment and / or size information of a size of the body part of the subject receiving treatment; displaying a second screen indicating connectivity between the device or an electronic device comprising the device and the treatment device; the operating state of the skin treatment device; a flash sound counter for counting the number of flash sounds detected since the start of the treatment operation or the predetermined or user-instructed moment; an indication of whether the flash sound counter has exceeded the flash sound threshold; and Indication that the treatment operation has ended to display a third screen indicating at least one of 8. Apparatus according to any one of claims 1 to 7, comprising an output for controlling a user interface.

11. a treatment device comprising one or more light sources that generate light pulses for application to the skin of a subject's body to perform a treatment action, the treatment device generating a flash sound for each light pulse generated by the treatment device; an audio sensor that detects an audio signal during the treatment operation; An apparatus according to any one of claims 1 to 10; A system comprising:

12. 1. A method of non-therapeutic photoepilation for use with a treatment device for performing a treatment action on a body part of a subject, the treatment device being configured to apply light pulses to skin of the body part to perform the treatment action, and to generate a flashing sound for each light pulse generated by the treatment device, the method comprising: acquiring an audio signal sensed during said treatment operation; segmenting said audio signal into audio segments, in particular audio segments of predetermined or pre-determined duration; extracting one or more characteristic features for each audio segment from the audio segments; detecting, for each audio segment, whether the audio segment includes a flash sound generated by the treatment device for each light pulse by applying a trained algorithm or computing system that has been trained with a plurality of audio segments, each audio segment including one or more of the characteristic features, to detect within the audio segment a flash sound generated by the treatment device for each light pulse; varying a flash sound counter that counts the number of flash sounds detected since the start of the treatment operation or a predetermined or user-instructed moment; said flash sound counter; an indication of whether the flash sound counter has exceeded a flash sound threshold; and Indication that the treatment operation has ended and outputting one or more of: A method comprising:

13. 1. A method of non-therapeutic photoepilation for use with a treatment device for performing a treatment action on a body part of a subject, the treatment device being configured to apply light pulses to skin of the body part to perform the treatment action, and to generate a flashing sound for each light pulse generated by the treatment device, the method comprising: acquiring a sensed or artificially produced audio signal during a treatment action; segmenting said audio signal into audio segments, in particular audio segments of predetermined or pre-determined duration; extracting one or more characteristic features for each audio segment from the audio segments; - deriving, for each audio segment, from annotations of the audio signal, flash sound information indicating whether the audio segment includes a flash sound generated by the treatment device for each light pulse, wherein the annotations of the audio signal indicate time information within the audio signal of the flash sound generated by the treatment device for each light pulse; providing the extracted one or more characteristic features and the derived flash sound information to an algorithm or computing system to train the algorithm or computing system to detect, for each audio segment, a flash sound generated by the treatment device for each light pulse within the audio segment based on the one or more characteristic features; A method comprising:

14. displaying a first screen requesting user input regarding the treatment operation; acquiring the user input via the first screen, the user input including an image of the body part of the subject receiving treatment and / or size information of the size of the body part of the subject receiving treatment; displaying a second screen indicating connectivity between the apparatus or electronic device comprising the user interface and the treatment device; the operating state of the skin treatment device; a flash sound counter for counting the number of flash sounds detected since the start of the treatment operation or the predetermined or user-instructed moment; the indication of whether the flash sound counter has exceeded the flash sound threshold; and the indication that the treatment operation has ended; to display a third screen indicating at least one of The method of claim 12, further comprising controlling a user interface.

15. A computer program comprising program code means for causing a computer to carry out the steps of the method according to claim 12, 13 or 14 when said computer program is executed on said computer.