Method, apparatus and program for improving recognition accuracy of acoustic data
The method enhances acoustic event recognition accuracy by frame-based post-processing correction, addressing noise interference and false alarms to improve sound detection reliability.
Patent Information
- Application Number
- JP2024526923
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-22
- Filing Date
- 2022-11-11
- Publication Date
- 2025-10-20
- Estimated Expiration
- 2042-11-11
AI Technical Summary
Existing acoustic event recognition technologies struggle with low accuracy in real-life environments due to ambient noise interference, difficulty in distinguishing between multiple events, and high false alarm rates, especially for hearing-impaired individuals and others with limited sound perception.
A method involving post-processing correction of acoustic data by constructing frames from time-series data, applying an acoustic recognition model, and performing threshold and time series analyses to transform and correct predicted values, reducing erroneous detections.
Improves the recognition accuracy of acoustic events by minimizing false alarms and enhancing the reliability of sound detection in noisy environments.
Smart Images

Figure 0007756381000001 
Figure 0007756381000002 
Figure 0007756381000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for improving the recognition rate of acoustic data, and more particularly to a technique for improving the recognition rate through post-processing correction of acoustic data. [Background technology]
[0002] Hearing-impaired people who cannot hear sounds at all or cannot distinguish sounds well have difficulty in daily life because they have difficulty hearing sounds and judging the situation. They also cannot use sound information to recognize dangerous situations in indoor or outdoor environments, making it impossible to respond immediately. In situations where hearing is limited or absent, such as hearing-impaired people, pedestrians wearing earphones, or the elderly, sounds occurring around the user may be blocked. Furthermore, in situations where it is difficult for users to detect sounds, such as when they are sleeping, they may encounter dangerous situations or be involved in accidents because they are unable to recognize the surrounding situation. Meanwhile, there is an increasing need for the development of technology to detect and recognize acoustic events in such environments. Technology to detect and recognize acoustic events is being continuously researched as it can be applied to various fields such as recognition of real-life context, recognition of dangerous situations, recognition of media content, and situation analysis in wired communications. The main research in acoustic event recognition technology is to extract various feature values such as MFCC, energy, spectral flux, and zero crossing rate from audio signals and verify their superior features, as well as research into Gaussian mixture model or rule-based classification methods.Recently, research has been conducted into deep learning-based machine learning methods to improve on these methods.However, these methods have limitations in that the accuracy of sound detection is guaranteed at low signal-to-noise ratios, making it difficult to distinguish between surrounding noise and the sound of an incident. That is, reliable acoustic event detection can be difficult in real-life environments containing a variety of ambient noises. Specifically, to detect valid acoustic events, it is necessary to determine whether an acoustic event has occurred from acoustic data acquired in a time series (i.e., continuously), and also to recognize what event class has occurred, making it difficult to ensure high reliability. Furthermore, when two or more events occur simultaneously, the recognition problem must be solved not only for a single event (monophonic) but also for multiple events (polyphonic), which can further reduce the acoustic event recognition rate. In addition, the reason why the recognition rate is low when detecting acoustic events from acoustic data acquired in real life is because there is a probability of false alarms, that is, the probability of determining that an event exists when no acoustic event has occurred, or determining that an event does not exist when an event has occurred.
[0003] Therefore, reducing the probability of erroneous detection in response to time-series acquired acoustic data may enable detection of acoustic events with increased confidence in real-life environments. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Republic of Korea Registered Patent 10-2014-0143069 Summary of the Invention [Problem to be solved by the invention]
[0005] The problem to be solved by the present invention is to solve the above-mentioned problems, and to provide an acoustic data recognition environment with improved accuracy through post-processing correction related to acoustic data. The problems to be solved by the present invention are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0006] To solve the above-mentioned problems, various embodiments of the present invention disclose a method for improving recognition accuracy of acoustic data, the method including the steps of: constructing one or more acoustic frames based on acoustic data, processing each of the one or more acoustic frames as an input of an acoustic recognition model to output a predicted value corresponding to each acoustic frame, identifying one or more recognition acoustic frames through a threshold value analysis based on the predicted value corresponding to each acoustic frame, identifying a transformed acoustic frame through a time series analysis based on the one or more recognition acoustic frames, and performing a transformation on the predicted value corresponding to the transformed acoustic frame.
[0007] In an alternative embodiment, the step of constructing one or more audio frames based on the audio data may include the step of dividing the audio data to have a predetermined first time unit size to construct the one or more audio frames. In an alternative embodiment, the start time of each of the one or more audio frames may be determined to have a difference of a magnitude of the second time unit from the start time of each of the adjacent audio frames. In an alternative embodiment, the predicted value may include one or more pieces of predicted item information and predicted numerical information corresponding to each of the one or more pieces of predicted item information, and the critical value analysis may be an analysis that identifies the one or more recognized sound frames by determining whether each of the one or more pieces of predicted numerical information corresponding to each of the sound frames is equal to or greater than a predetermined critical value corresponding to each of the predicted item information. In an alternative embodiment, the step of identifying the transformed audio frame through the time series analysis may include the steps of identifying predicted item information corresponding to each of the one or more recognition audio frames, determining whether the identified predicted item information is repeated a predetermined number of times or more during a predetermined reference time, and identifying the transformed audio frame based on the determination result. In an alternative embodiment, the method may include a step of identifying a correlation between each of the recognition audio frames based on prediction item information corresponding to each of the one or more recognition audio frames, and a step of determining whether or not to adjust a critical value and a critical number of times corresponding to each of the one or more audio frames based on the correlation. In an alternative embodiment, the transformation of the predicted value may include at least one of a noise transformation that transforms the output of the acoustic recognition model based on the transformed acoustic frame into an unrecognized item, and an acoustic item transformation that transforms predicted item information associated with the transformed acoustic frame into calibrated predicted item information. In an alternative embodiment, the proofreading prediction item information may be determined based on a correlation between the prediction item information.
[0008] According to another embodiment of the present invention, there is disclosed an apparatus for performing a method for improving recognition accuracy of acoustic data, the apparatus including a memory for storing one or more instructions and a processor for executing the one or more instructions stored in the memory, the processor executing the one or more instructions to perform the method for improving recognition accuracy of acoustic data. According to yet another embodiment of the present invention, there is disclosed a computer program stored on a computer-readable recording medium, which can be coupled to a computer as hardware to perform the method for improving the recognition accuracy of acoustic data described above. Other specific details of the invention are included in the detailed description and drawings. [Effects of the Invention]
[0009] Various embodiments of the present invention can provide the effect of improving the recognition accuracy of sound data through correction of sound data. The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating a system for performing a method for improving recognition accuracy of acoustic data according to an embodiment of the present invention. [Figure 2] 1 is a hardware configuration diagram of a server for improving the recognition accuracy of acoustic data according to an embodiment of the present invention; [Figure 3] 1 shows a flow chart illustrating an exemplary method for improving recognition accuracy of acoustic data in accordance with one embodiment of the present invention. [Figure 4] 1 illustrates an example diagram for explaining a process of constructing one or more acoustic frames based on acoustic data according to an embodiment of the present invention; [Figure 5] 1 is a diagram illustrating an example of a process in which an acoustic recognition model according to an embodiment of the present invention outputs a predicted value based on an acoustic frame; [Figure 6] 1 is a flowchart illustrating an exemplary critical value analysis process related to one embodiment of the present invention. [Figure 7] 1 shows a flowchart illustrating an exemplary time series analysis process associated with one embodiment of the present invention. [Figure 8] 1 illustrates an exemplary table for explaining the acoustic data correction process associated with one embodiment of the present invention. [Figure 9] 1 is a diagram illustrating an example of a process for correcting acoustic data according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0011] Various embodiments will now be described with reference to the drawings. Various details are provided herein to provide an understanding of the present invention. However, it will be apparent that such embodiments may be practiced without such specific details. As used herein, the terms "component," "module," "system," and the like refer to computer-related entities, hardware, firmware, software, a combination of software and hardware, or the execution of software. For example, a component may be, but is not limited to, a procedure running on a processor, a processor, an object, a thread of execution, a program, and / or a computer. For example, an application running on a computing device and the computing device may both be a component. One or more components may reside within a processor and / or thread of execution. A component may be localized within one computer. A component may be distributed among two or more computers. Such components may also execute from various computer-readable media having various data structures stored therein. Components may communicate via local and / or remote processes, for example, via signals comprising one or more data packets (e.g., data from one component interacting with other components in a local system, a distributed system, and / or data transmitted over a network such as the Internet to other systems via signals). Also, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X utilizes A or B" is intended to mean one of the natural inclusive permutations. That is, if X utilizes A; X utilizes B; or X utilizes both A and B, then "X utilizes A or B" can apply to any of these cases. Also, as used herein, the term "and / or" should be understood to refer to and include all possible combinations of one or more of the associated listed items. Additionally, the terms "comprise" and / or "comprises" should be understood to mean the presence of the relevant feature and / or component. However, the terms "comprise" and / or "comprises" should be understood not to exclude the presence or addition of one or more other features, components and / or groups thereof. Furthermore, unless otherwise specified or clear from the context as referring to the singular form, the singular in the present specification and claims should generally be interpreted to mean "one or more."
[0012] Those skilled in the art should additionally recognize that the various illustrative logical blocks, components, modules, circuits, means, logic, and algorithm steps described in connection with the embodiments disclosed herein may be embodied in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the various illustrative components, blocks, components, means, logic, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is embodied in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in various ways for each particular application. However, such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0013] The description of the illustrated embodiments is provided to enable one of ordinary skill in the art to make and practice the invention. Various modifications to these embodiments will be apparent to those of ordinary skill in the art. The generic principles defined herein may be applied to other embodiments without departing from the scope of the invention, and the invention is not limited to the embodiments shown herein. The invention is to be accorded the widest scope consistent with the principles and novel features disclosed herein. In this specification, the term "computer" refers to any type of hardware device including at least one processor, and may also encompass software configurations operating on the hardware device, depending on the embodiment. For example, the term "computer" may be understood to include, but is not limited to, smartphones, tablet PCs, desktops, laptops, and user clients and applications running on each device.
[0014] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. Although each step described in this specification is described as being performed by a computer, the subject of each step is not limited to this, and depending on the embodiment, at least some of each step may be performed by different devices. Here, a method for improving the recognition accuracy of acoustic data according to various embodiments of the present invention may relate to a method for correcting acoustic data so as to improve the recognition rate of the acoustic data. Correction of acoustic data may mean, for example, post-processing correction related to the acoustic data. That is, when time-series acoustic data is acquired, the present invention may improve the accuracy in the recognition process of the acoustic data by performing post-processing correction on the corresponding acoustic data. In the embodiment, improvement of recognition accuracy of acoustic data may mean improvement of recognition accuracy of detecting a specific event in the acoustic data.
[0015] Meanwhile, in order to detect or recognize a specific event in acoustic data with high accuracy, it may be important to reduce the probability of erroneous detection, which may refer to the probability of determining that an acoustic event exists when it does not occur, or determining that an event does not exist when it does occur. According to an embodiment, a method for improving the accuracy of audio data recognition may minimize the probability of error detection of audio data by dividing audio data into a plurality of audio frames each having a certain time unit, and then performing audio recognition for each of the divided audio frames to improve the accuracy of audio data recognition. In this case, each audio frame may have at least a partial overlap with other audio frames. That is, the present invention may subdivide time-series audio data into a plurality of audio frames, analyze each audio frame, and perform transformation on at least some of the audio frames as a result of the analysis. For example, it may be determined that a specific audio has been recognized only if the specific audio (e.g., a siren sound) is recognized across at least two audio frames among the plurality of audio frames. In other words, if the specific audio (e.g., a siren sound) is recognized only in a specific audio frame among the plurality of audio frames (i.e., if the specific audio is not recognized in audio frames adjacent to the specific audio frame), it may be determined that the specific audio has not been recognized, and transformation related to the corresponding audio frame may be performed. Here, conversion related to an audio frame may mean, for example, converting a sound recognized in relation to a specific audio frame (e.g., a siren sound) into an unrecognized sound, or converting the sound into another sound (e.g., a sound not related to the recognition), since the sound is an erroneous recognition sound. In other words, a sound recognized in only one frame can be eliminated as an error, and a sound recognized continuously in relation to the frame can be determined as a correctly recognized sound.
[0016] In summary, the present invention can improve the recognition accuracy of the entire sound data by dividing the sound data into frames and determining that sounds that are not continuously recognized in each frame are misrecognized sounds and performing post-processing correction. A more detailed description of the method for improving the recognition accuracy of sound data will be provided below.
[0017] Figure 1 is a diagram illustrating a system for performing a method for improving recognition accuracy of acoustic data according to an embodiment of the present invention. As shown in Figure 1, the system for performing a method for improving recognition accuracy of acoustic data according to an embodiment of the present invention may include a server 100 for improving recognition accuracy of acoustic data, a user terminal 200, and an external server 300. Here, the system for performing a method for improving recognition accuracy of acoustic data illustrated in Figure 1 is according to one embodiment, and its components are not limited to those illustrated in Figure 1, and may be added, changed, or deleted as necessary.
[0018] In one embodiment, the server 100 for improving the recognition accuracy of sound data may determine whether a specific event has occurred based on the sound data. Specifically, the server 100 for improving the recognition accuracy of sound data may acquire real-life sound data and determine whether a specific event has occurred by analyzing the acquired sound data. In one embodiment, the specific event may be related to the occurrence of security, safety, or danger, such as an alarm sound, a child crying, breaking glass, or a flat tire. The specific description of the sound related to the specific event described above is merely an example, and the present invention is not limited thereto.
[0019] According to an embodiment, since sound data acquired in real life contains various ambient noises, it may be difficult to detect sound events with high reliability. Accordingly, when the server 100 for improving the recognition accuracy of sound data receives sound data, the server 100 for improving the recognition accuracy of sound data of the present invention may perform post-processing correction on the sound data. Here, the post-processing correction may refer to correction for reducing the probability of error detection during the sound data recognition process. For example, the post-processing correction may include converting a recognized sound (e.g., the sound of glass breaking) in a certain section of the sound data into an unrecognized sound (i.e., processing it as noise) or converting the recognized sound into a different sound. In other words, the server 100 for improving the recognition accuracy of sound data may acquire time-series sound data related to real life and ensure improved recognition accuracy through post-processing correction on the acquired sound data. According to an embodiment, the server 100 for improving the recognition accuracy of acoustic data may include any server implemented by an API (Application Programming Interface). For example, the user terminal 200 may acquire acoustic data and transmit it to the server 100 through the API. For example, the server 100 may acquire acoustic data from the user terminal 200 and determine that an emergency alarm sound (e.g., a siren sound) has been generated through analysis of the acoustic data. In an embodiment, the server 100 for improving the recognition accuracy of acoustic data may analyze the acoustic data through an acoustic recognition model (e.g., an artificial intelligence model).
[0020] In one embodiment, an acoustic recognition model (e.g., an artificial intelligence model) is comprised of one or more network functions, which may be comprised of a collection of interconnected computational units that may generally be referred to as "nodes." Such "nodes" may also be referred to as "neurons." One or more network functions are comprised of at least one or more nodes. The nodes (or neurons) that comprise one or more network functions may be interconnected by one or more "links." Within an artificial intelligence model, one or more nodes connected through links can form a relative input node-output node relationship. The concepts of input node and output node are relative, and any node that has an output node relationship with one node can also have an input node relationship with another node, and vice versa. As mentioned above, the input node-output node relationship can be generated around links. One or more output nodes can be connected to one input node through links, and vice versa. In a relationship between an input node and an output node connected through a link, the value of the output node may be determined based on data input to the input node. Here, the node connecting the input node and the output node may have a weight. The weight may be variable and may be changed by a user or an algorithm so that the artificial intelligence model performs a desired function. For example, when one or more input nodes are connected to one output node through respective links, the output node may determine its output node value based on the value input to the input node connected to the output node and the weight assigned to the link corresponding to each input node.
[0021] As described above, an AI model has one or more nodes interconnected through one or more links to form a relationship between input nodes and output nodes within the AI model. The characteristics of the AI model can be determined by the number of nodes and links, the correlation between the nodes and links, and the weights assigned to each link within the AI model. For example, if two AI models have the same number of nodes and links but different weights between the links, the two AI models can be recognized as different from each other. Some of the nodes constituting an AI model may constitute a layer based on their distance from the first input node. For example, a set of nodes whose distance from the first input node is n may constitute n layers. The distance from the first input node may be defined by the minimum number of links that must be traversed to reach the corresponding node from the first input node. However, this definition of a layer is arbitrary for the purpose of explanation, and the order of layers within an AI model may be defined in a manner different from that described above. For example, the layer of a node may be defined by its distance from the final output node.
[0022] The first input node may refer to one or more nodes to which data is directly input without a link in relation to other nodes within the artificial intelligence model. Alternatively, it may refer to a node that does not have any other input nodes connected by a link in relation between nodes based on links within the artificial intelligence model network. Similarly, the final output node may refer to one or more nodes that do not have any output nodes in relation to other nodes within the artificial intelligence model. Furthermore, the hidden node may refer to a node that constitutes the artificial intelligence model, rather than the first input node and the final output node. The artificial intelligence model according to an embodiment of the present invention may have more nodes in the input layer than in the hidden layer closer to the output layer, and may be an artificial intelligence model in which the number of nodes decreases as one progresses from the input layer to the hidden layer.
[0023] An artificial intelligence model can include one or more hidden layers. Hidden nodes in a hidden layer can receive inputs from the outputs of previous layers and surrounding hidden nodes. The number of hidden nodes in each hidden layer can be the same or different. The number of nodes in the input layer can be determined based on the number of data fields in the input data and can be the same or different from the number of hidden nodes. Input data input to the input layer can be operated on by hidden nodes in the hidden layer and output by a fully connected layer (FCL), which is an output layer.
[0024] In various embodiments, the AI model may be subjected to supervised learning using a plurality of acoustic data and feature information corresponding to each acoustic data as training data, but is not limited thereto and various learning methods may be applied. Here, supervised learning is a method of generating learning data by labeling specific data and information related to the specific data, and using the generated learning data for learning. It also refers to a method of generating learning data by labeling two pieces of data that have a causal relationship, and learning through the generated learning data.
[0025] In one embodiment, the server 100 for improving the recognition accuracy of acoustic data may determine whether to stop the training using verification data when the training of one or more network functions has been performed for a predetermined number of epochs or more. The predetermined epochs may be a portion of the overall training target epoch. The verification data may be composed of at least a portion of the labeled training data. That is, the server 100 for improving the recognition accuracy of acoustic data performs training of an AI model using the training data, and after the training of the AI model is repeated for a predetermined number of epochs or more, the server 100 may determine whether the learning effect of the AI model is at or above a predetermined level using the verification data. For example, when the server 100 for improving the recognition accuracy of acoustic data performs training with a target number of iterative training times of 10 using 100 pieces of training data, the server 100 may perform 10 iterative training times, which is the predetermined epoch, and then perform three iterative training times using 10 pieces of verification data. If the change in the AI model output during the three iterative training times is below a predetermined level, the server 100 may determine that further training is meaningless and terminate the training.
[0026] That is, the validation data may be used to determine the completion of learning based on whether the effect of each epoch of the iterative learning of the AI model is above or below a certain level. The numbers of learning data, validation data, and the number of iterations described above are merely examples, and the present invention is not limited thereto. The server 100 for improving the recognition accuracy of acoustic data may generate an AI model by testing the performance of one or more network functions using test data and determining whether to activate one or more network functions. The test data may be used to verify the performance of the AI model and may comprise at least a portion of the training data. For example, 70% of the training data may be used for training the AI model (i.e., training to adjust weights to output result values similar to the labels), and 30% may be used as test data for verifying the performance of the AI model. The server 100 for improving the recognition accuracy of acoustic data may input the test data into the AI model that has completed training, measure the error, and determine whether to activate the AI model based on whether the error is above a predetermined performance level.
[0027] The server 100 for improving the recognition accuracy of acoustic data verifies the performance of the AI model that has completed training using test data for the AI model, and if the performance of the AI model that has completed training is above a predetermined standard, the AI model can be activated for use in other applications. In addition, the server 100 for improving the recognition accuracy of acoustic data may deactivate and discard an AI model that has completed training if the performance of the model is below a predetermined standard. For example, the server 100 for improving the recognition accuracy of acoustic data may determine the performance of the generated AI model based on factors such as accuracy, precision, and recall. The performance evaluation criteria described above are merely exemplary, and the present invention is not limited thereto. The server 100 for improving the recognition accuracy of acoustic data may generate multiple AI models by training each AI model independently, evaluate their performance, and use only AI models that meet or exceed a certain level of performance. However, the present invention is not limited thereto.
[0028] Throughout this specification, the terms computational model, neural network, network function, and neural network may be used interchangeably (hereinafter, they will be referred to as a neural network). A data structure may include a neural network. A data structure including a neural network may be stored on a computer-readable medium. A data structure including a neural network may also include data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and a loss function for training the neural network. A data structure including a neural network may include any of the components disclosed above. That is, a data structure including a neural network may include all or any combination of data input to the neural network, neural network weights, neural network hyperparameters, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and a loss function for training the neural network. In addition to the above-mentioned components, a data structure including a neural network may include any other information that determines the characteristics of a neural network. Furthermore, a data structure may include any type of data used or generated in the computational process of a neural network, and is not limited to the above. The computer-readable medium may include a computer-readable recording medium and / or a computer-readable transmission medium. A neural network may be composed of a collection of interconnected computational units, which may generally be referred to as nodes. Such nodes may also be referred to as neurons. A neural network is composed of at least one or more nodes.
[0029] According to an embodiment of the present invention, the server 100 for improving the recognition accuracy of acoustic data may be a server that provides a cloud computing service. More specifically, the server 100 for improving the recognition accuracy of acoustic data may be a server that provides a cloud computing service, which is a type of Internet-based computing, in which information is processed not on a user's computer but on another computer connected to the Internet. The cloud computing service may store data on the Internet and allow users to access the data or programs they need anytime and anywhere through an Internet connection without having to install them on their own computers. The cloud computing service may also allow users to easily share and transmit data stored on the Internet with simple operations and clicks. In addition, the cloud computing service may not only simply store data on an Internet server, but may also allow users to perform desired tasks using the functions of application programs provided on the web without installing additional programs, and may allow multiple users to work while sharing documents simultaneously. The cloud computing service may be implemented in at least one form of Infrastructure as a Service (IaaS), Platform as a Service (PaaS), Software as a Service (SaaS), a virtual machine-based cloud server, or a container-based cloud server. That is, the server 100 for improving the recognition accuracy of acoustic data according to the present invention may be implemented in the form of at least one of the cloud computing services described above. The specific description of the cloud computing service described above is merely an example, and the present invention may include any platform that establishes a cloud computing environment.
[0030] In various embodiments, the server 100 for improving the recognition accuracy of acoustic data can be connected to the user terminal 200 via a network, and can not only generate and provide an acoustic recognition model for analyzing acoustic data, but also provide information (e.g., acoustic event information) analyzed from the acoustic data through the acoustic recognition model to the user terminal. Here, a network can refer to a connection structure that allows information exchange between nodes such as multiple terminals and servers. For example, networks include local area networks (LANs), wide area networks (WANs), the Internet (WWW), wired / wireless data communication networks, telephone networks, and wired / wireless television communication networks.
[0031] In addition, wireless data communication networks include, but are not limited to, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), 5GPP (5th Generation Partnership Project), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, the Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), RF (Radio Frequency), Bluetooth network, NFC (Near-Field Communication) network, satellite broadcasting network, analog broadcasting network, DMB (Digital Multimedia Broadcasting) network, etc.
[0032] In one embodiment, the user terminal 200 may be connected to a server 100 for improving the accuracy of recognition of sound data via a network, and may provide sound data to the server 100 for improving the accuracy of recognition of sound data, and may receive information related to the occurrence of various events (e.g., the occurrence of an alarm sound, a child crying, the sound of glass breaking, the sound of a flat tire, etc.) in response to the provided sound data. Here, the user terminal 200 is a wireless communication device that ensures portability and mobility, and may include, but is not limited to, all kinds of handheld-based wireless communication devices such as navigation systems, Personal Communication Systems (PCS), Global System for Mobile communications (GSM), Personal Digital Cellular (PDC), Personal Handyphone Systems (PHS), Personal Digital Assistants (PDAs), International Mobile Telecommunication (IMT)-2000, Code Division Multiple Access (CDMA)-2000, W-Code Division Multiple Access (W-CDMA), Wireless Broadband Internet (Wibro) terminals, smartphones, smartpads, tablet PCs, etc. For example, the user terminal 200 may be installed in a specific area to perform sensing related to the specific area. For example, the user terminal 200 may be installed in a vehicle and acquire acoustic data generated while the vehicle is parked or running. The above description of the specific location or place where the user terminal is provided is merely an example, and the present invention is not limited thereto.
[0033] In one embodiment, the external server 300 may be connected to the server 100 for improving the recognition accuracy of acoustic data via a network, and may provide various information / data necessary for the server 100 for improving the recognition accuracy of acoustic data to analyze acoustic data using an artificial intelligence model, or may receive, store, and manage result data derived from analyzing acoustic data using the artificial intelligence model. For example, the external server 300 may be, but is not limited to, a storage server separately installed outside the server 100 for improving the recognition accuracy of acoustic data. Hereinafter, a hardware configuration of the server 100 for improving the recognition accuracy of acoustic data will be described with reference to FIG. 2.
[0034] FIG. 2 is a hardware configuration diagram of a server for improving the recognition accuracy of acoustic data according to an embodiment of the present invention. 2, a server 100 for improving recognition accuracy of acoustic data according to an embodiment of the present invention (hereinafter referred to as "server 100") may include one or more processors 110, a memory 120 for loading a computer program 151 executed by the processor 110, a bus 130, a communication interface 140, and a storage 150 for storing the computer program 151. Only components relevant to the embodiment of the present invention are shown in FIG. 2. Therefore, a person skilled in the art to which the present invention pertains will understand that other general-purpose components may be included in addition to the components shown in FIG. 2. The processor 110 controls the overall operation of each component of the server 100. The processor 110 may be configured to include a central processing unit (CPU), a microprocessor unit (MPU), a microcontroller unit (MCU), a graphic processing unit (GPU), or any other type of processor widely known in the technical field of the present invention. The processor 110 can read a computer program stored in the memory 120 and perform data processing for an artificial intelligence model according to an embodiment of the present invention. According to an embodiment of the present invention, the processor 110 can perform calculations for neural network training. The processor 110 can perform calculations for neural network training, such as processing input data for training using deep learning (DL), extracting features from the input data, calculating errors, and updating the weights of the neural network using backpropagation.
[0035] In addition, at least one of the processors 110, a CPU, a GPGPU, and a TPU, may process network function learning. For example, the CPU and the GPGPU may both process network function learning and data classification using the network function. In addition, in one embodiment of the present invention, processors of multiple computing devices may be used together to process network function learning and data classification using the network function. In addition, a computer program executed in a computing device according to an embodiment of the present invention may be a CPU, GPGPU, or TPU executable program. As used herein, the term "network function" may be used interchangeably with "artificial neural network" or "neural network." As used herein, the term "network function" may include one or more neural networks, in which case the output of the network function may be an ensemble of the outputs of the one or more neural networks.
[0036] The processor 110 can provide an acoustic recognition model according to an embodiment of the present invention by reading a computer program stored in the memory 120. According to an embodiment of the present invention, the processor 110 can perform calculations for training the acoustic recognition model. According to one embodiment of the present invention, the processor 110 may generally process the overall operation of the server 100. The processor 110 may process signals, data, information, etc. input or output through the components detailed above, or may run application programs stored in the memory 120, thereby providing or processing appropriate information or functions to a user or a user terminal.
[0037] Furthermore, the processor 110 can execute operations for at least one application or program for executing a method according to an embodiment of the present invention, and the server 100 can include one or more processors. In various embodiments, the processor 110 may further include a random access memory (RAM, not shown) and a read-only memory (ROM, not shown) that temporarily and / or permanently store signals (or data) processed within the processor 110. The processor 110 may also be implemented in the form of a system on a chip (SoC) that includes at least one of a graphics processing unit, RAM, and ROM. The memory 120 stores various data, instructions, and / or information. The memory 120 can load a computer program 151 from the storage 150 to perform the methods / operations according to various embodiments of the present invention. When the computer program 151 is loaded into the memory 120, the processor 110 can perform the methods / operations by executing one or more instructions constituting the computer program 151. The memory 120 can be implemented as a volatile memory such as a RAM, although the scope of the present invention is not limited thereto.
[0038] The bus 130 provides a communication function between the components of the server 100. The bus 130 may be implemented as various types of buses such as an address bus, a data bus, and a control bus. The communication interface 140 supports wired / wireless Internet communication of the server 100. The communication interface 140 may also support various communication methods other than Internet communication. To this end, the communication interface 140 may be configured to include a communication module that is well known in the art of the present invention. In some embodiments, the communication interface 140 may be omitted.
[0039] The storage 150 may non-temporarily store a computer program 151. When a process for improving the recognition accuracy of acoustic data is performed through the server 100, the storage 150 may store various information required to provide the process for improving the recognition accuracy of acoustic data. Storage 150 may be configured to include non-volatile memory such as ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), flash memory, etc., a hard disk, a removable disk, or any form of computer-readable recording medium widely known in the technical field to which the present invention belongs. The computer program 151 may include one or more instructions that, when loaded into the memory 120, cause the processor 110 to perform the methods / operations according to various embodiments of the present invention. That is, the processor 110 can perform the methods / operations according to various embodiments of the present invention by executing the one or more instructions.
[0040] In one embodiment, the computer program 151 may include one or more instructions for performing a method for improving the recognition accuracy of acoustic data, including the steps of: constructing one or more acoustic frames based on acoustic data; processing each of the one or more acoustic frames as an input for an acoustic recognition model to output a predicted value corresponding to each acoustic frame; identifying one or more recognition acoustic frames through a threshold value analysis based on the predicted value corresponding to each acoustic frame; identifying a transformed acoustic frame through a time series analysis based on the one or more recognition acoustic frames; and performing a transformation on the predicted value corresponding to the transformed acoustic frame.
[0041] The steps of a method or algorithm described in connection with the embodiments of the present invention may be embodied directly in hardware, in a software module executed by hardware, or in a combination thereof. The software module may reside in Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), Flash Memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium commonly known in the art to which the present invention pertains. The components of the present invention may be embodied as a program (or application) stored on a medium for execution in combination with a computer, which is hardware. The components of the present invention may be implemented as software programming or software elements. Similarly, embodiments include various algorithms embodied as a combination of data structures, processes, routines, or other programming constructs, and may be implemented in programming or scripting languages such as C, C++, Java, assembler, etc. Functional aspects may be embodied as algorithms executed by one or more processors. A method for improving the recognition accuracy of acoustic data, performed by server 100, will now be described with reference to FIGS. 3 to 9.
[0042] 3 is a flowchart illustrating an example of a method for improving recognition accuracy of acoustic data according to an embodiment of the present invention. The order of the steps illustrated in FIG. 3 may be changed as necessary, and at least one step may be omitted or added. That is, the following steps are merely an embodiment of the present invention, and the scope of the present invention is not limited thereto. According to an embodiment of the present invention, the server 100 may acquire acoustic data. The acoustic data may include information related to sounds acquired in real life. Acquiring acoustic data according to an embodiment of the present invention may involve receiving or loading acoustic data stored in the memory 120. Acquiring acoustic data may also involve receiving or loading data from another storage medium, another computing device, or a separate processing module within the same computing device via wired or wireless communication.
[0043] According to one embodiment, the acoustic data may be acquired through a user terminal 200 associated with a user. For example, the user terminal 200 associated with a user may include any kind of handheld-based wireless communication device such as a smartphone, a smartpad, a tablet PC, etc., or an electronic device (e.g., a device capable of receiving acoustic data through a microphone) provided in a specific space (e.g., a user's living space). According to one embodiment of the present invention, the server 100 may generate one or more audio frames based on audio data (S100). The one or more audio frames may be obtained by dividing audio data, which is time-series information, into a plurality of frames based on a specific time unit. Specifically, the server 100 may generate one or more audio frames by dividing the audio data to have a predetermined first time unit size. For example, if the first audio data is audio data acquired corresponding to a time period of one minute, the server 100 may generate 30 audio frames by setting the first time unit to two seconds. The specific numerical values related to the first time unit and the one or more audio frames described above are merely examples, and the present invention is not limited thereto.
[0044] According to one embodiment, the server 100 may configure one or more audio frames so that each of the one or more audio frames at least partially overlaps with the other. Referring to FIG. 4, the start point of each of the one or more audio frames may be determined to have the same size as the start point of each adjacent audio frame (second time unit 400b). According to one embodiment, the size of the second time unit 400b may be determined to be smaller than the size of the first time unit 400a. That is, as shown in FIG. 4, the server 100 may generate one or more audio frames (i.e., first audio frame 411, second audio frame 412, third audio frame 413, etc.) 410 having the same first time unit 400a. In this case, each audio frame may be configured to differ from each adjacent audio frame by the size of the second time unit 400b, which is smaller than the size of the first time unit 400a. Accordingly, each audio frame may at least partially overlap with each adjacent audio frame. As a specific example, if the acoustic data 400 relates to acoustic data acquired over a 10-second period, the first time unit 400a may be set to 2 seconds, and the second time unit 400b may be set to 1 second, which is shorter than the first time unit 400a. In this case, the first acoustic frame 411 may relate to acoustic data acquired over a period of 0 to 2 seconds, the second acoustic frame 412 may relate to acoustic data acquired over a period of 1 to 3 seconds, and the third acoustic frame 413 may relate to acoustic data acquired over a period of 2 to 4 seconds. The specific numerical values relating to the total time, first time unit, and second time unit of the acoustic data described above are merely examples, and the present invention is not limited thereto. That is, one or more acoustic frames are configured so that the start point of each acoustic frame has a difference in size between the start point of each adjacent acoustic frame and the start point of the second time unit 400b that is smaller than the size of the first time unit 400a, so that at least a portion of each acoustic frame can have an overlapping section.
[0045] According to one embodiment of the present invention, the server 100 processes each of one or more audio frames as an input of an audio recognition model and outputs a predicted value corresponding to each audio frame (S200). According to one embodiment, the server 100 can train an autoencoder through unsupervised learning. Specifically, the server 100 can train a dimension reduction network function (e.g., an encoder) and a dimension restoration network function (e.g., a decoder) that configure the autoencoder to output output data similar to the input data. In more detail, the dimension reduction network function can learn only core feature data (or features) of the input audio data through a hidden layer during the encoding process, while losing the remaining information. In this case, the output data of the hidden layer during the decoding process through the dimension restoration network function can be an approximation of the input data (i.e., audio data) rather than a perfect copy. In other words, the server 100 can train the autoencoder by adjusting weights so that the output data and the input data are as similar as possible.
[0046] An autoencoder may be a type of neural network that outputs output data similar to input data. An autoencoder may include at least one hidden layer, and an odd number of hidden layers may be arranged between the input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called a bottleneck layer (encoding), and then expanded symmetrically from the bottleneck layer to the output layer (symmetric to the input layer). The number of input and output layers may correspond to the number of input data items remaining after preprocessing of the input data. In an autoencoder structure, the number of nodes in the hidden layer included in the encoder may decrease as it moves away from the input layer. The number of nodes in the bottleneck layer (the layer with the fewest nodes located between the encoder and decoder) may be maintained at a certain number or more (e.g., more than half of the input layer) because a small number of nodes may not transmit a sufficient amount of information.
[0047] The server 100 may input a training data set including a plurality of training data tagged with object information to a trained dimension reduction network, and match the output object-specific feature data with the tagged object information and store the matched object information. Specifically, the server 100 may use a dimension reduction network function to input a first training data subset tagged with first sound identification information (e.g., the sound of glass breaking) to acquire feature data of the first object for the training data included in the first training data subset. The acquired feature data may be expressed as a vector. In this case, the feature data output corresponding to each of the plurality of training data included in the first training data subset is output through the training data related to the first sound, and therefore may be located relatively close to each other in vector space. The server 100 may match the first sound identification information (i.e., the sound of glass breaking) with the feature data related to the first sound represented as a vector and store the matched first sound identification information. In the case of a trained autoencoder dimensionality reduction network function, the dimensionality restoration network function can be trained to extract features well that allow the input data to be well restored.
[0048] Furthermore, for example, the plurality of training data included in each first training data subset tagged with second sound discrimination information (e.g., the sound of a siren) may be converted into feature data (i.e., features) through a dimensionality reduction network function and displayed in a vector space. In this case, since the feature data is output through the training data related to the second sound discrimination information (i.e., the sound of a siren), it may be located relatively close to each other in the vector space. In this case, the feature data corresponding to the second sound discrimination information may be displayed in a different vector space from the feature data corresponding to the first sound discrimination information (e.g., the sound of glass breaking).
[0049] In an embodiment, the server 100 may configure the acoustic recognition model 500 by including a dimension reduction network function in the trained autoencoder. That is, when an acoustic frame is input, the acoustic recognition model 500 configured by including the dimension reduction network function generated through the above-described training process may extract feature information (i.e., features) corresponding to the acoustic frame through an operation using the dimension reduction network function from the corresponding acoustic frame. In this case, the acoustic recognition model 500 may evaluate the similarity of the acoustic style by comparing the distance in the vector space between the area where the feature corresponding to the acoustic frame is displayed and the object-specific feature data, and may output a predicted value corresponding to the acoustic data based on the similarity evaluation. In one embodiment, the predicted value may include one or more pieces of predicted item information and predicted numerical information corresponding to each of the one or more pieces of predicted item information.
[0050] Specifically, the acoustic recognition model 500 can output feature information (i.e., features) by operating an acoustic frame using a dimension reduction network function. In this case, the acoustic recognition model can include one or more predicted item information corresponding to the acoustic frame and predicted numerical information corresponding to each predicted item information based on the position between the feature information output corresponding to the acoustic frame and feature data for each acoustic identification information previously recorded in a vector space through learning. The one or more pieces of predicted item information are information about the type of sound associated with the sound, and may include, for example, the sound of glass breaking, a flat tire, an emergency siren activating, a puppy barking, or rain falling. Such predicted item information may be generated based on acoustic identification information located close to feature information output corresponding to an audio frame in vector space. For example, an audio recognition model may generate one or more pieces of predicted item information through acoustic identification information matched with feature information located close to first feature information output corresponding to a first audio frame. The specific description of the one or more pieces of predicted item information described above is merely illustrative, and the present invention is not limited thereto.
[0051] The predicted numerical information corresponding to each prediction item information may be information about a value predicted corresponding to each prediction item information. For example, the acoustic recognition model may configure one or more prediction item information through acoustic identification information matched with feature information located close to the first feature information output corresponding to the first acoustic frame. In this case, the closer the first feature information is to the feature information corresponding to each acoustic identification information, the higher the predicted numerical information may be output, and the farther the first feature information is from the feature information corresponding to each acoustic identification information, the lower the predicted numerical information may be output. 5, the acoustic recognition model 500 may output predicted item information 610 associated with a "siren sound," a "scream," a "glass breaking sound," and "other sounds" in response to a first acoustic frame 411. The acoustic recognition model 500 may also output predicted numerical information 620 of "1," "95," "3," and "2" in response to each piece of predicted item information 610. That is, the acoustic recognition model 500 may output predicted values 600 in which the probability associated with a siren sound is 1, the probability associated with a scream is 95, the probability associated with glass breaking is "3," and the probability associated with other sounds is "2" in response to the first acoustic frame 411. The specific numerical values for each of the predicted item information and predicted numerical information described above are merely examples, and the present invention is not limited thereto. That is, the server 100 can output a predicted value corresponding to one or more acoustic frames constructed based on acoustic data through the acoustic recognition model 500. For example, the acoustic recognition model 500 can output a first predicted value corresponding to the first acoustic frame 411, a second predicted value corresponding to the second acoustic frame 412, and a third predicted value corresponding to the third acoustic frame 413.
[0052] According to one embodiment of the present invention, the server 100 can identify one or more recognition audio frames through a threshold value analysis based on a prediction value corresponding to each audio frame (S300). Here, the threshold value analysis may refer to an analysis that identifies one or more recognition audio frames by determining whether one or more pieces of prediction numerical information corresponding to each audio frame are equal to or greater than a predetermined threshold value corresponding to each prediction item information. A detailed description of a method for identifying one or more recognition audio frames through the threshold value analysis will be provided below with reference to FIG. 6. In one embodiment, the server 100 may identify one or more pieces of predicted numerical information corresponding to one or more acoustic frames (S310). As a specific example, the one or more acoustic frames may include a first acoustic frame and a second acoustic frame. The server 100 may process each acoustic frame as an input to the acoustic recognition model 500 and output a predicted value corresponding to each acoustic frame. Here, the predicted value may include one or more pieces of predicted item information and predicted numerical information corresponding to each prediction item information. Accordingly, the server 100 may identify predicted numerical information corresponding to each acoustic frame through the predicted value output by the acoustic recognition model corresponding to each acoustic frame.
[0053] For example, the server 100 can identify that the predicted numerical information corresponding to the "sound of breaking glass" and the "crying of a child" are "82" and "5" respectively based on the predicted values corresponding to the first sound frame 411. For example, the server 100 may identify that the predicted numerical information corresponding to the "sound of breaking glass" and the "sound of a siren" are "50" and "12," respectively, based on the predicted values corresponding to the second sound frame 412. The specific numerical values for the predicted numerical information described above are merely examples, and the present invention is not limited thereto. The server 100 may also identify a predetermined threshold value corresponding to one or more pieces of prediction item information (S320). In one embodiment, a threshold value may be preset corresponding to each piece of prediction item information. The threshold value may refer to a threshold value for identifying an audio recognition result having a certain level of accuracy or higher. For example, if the prediction numerical information corresponding to a first audio frame is equal to or greater than the threshold value, it may indicate that the audio recognition result of the first audio frame is at a reliable level. As another example, if the prediction numerical information corresponding to a second audio frame is less than the threshold value, it may indicate that the audio recognition result of the second audio frame is somewhat less accurate. The specific descriptions of each audio frame above are merely examples, and the present invention is not limited thereto.
[0054] The threshold values may be set differently for each prediction item. According to an embodiment, the threshold values for each prediction item may be predetermined according to the difficulty of sound recognition. For example, the more difficult a sound is to recognize, the lower the threshold value may be set, and the easier a sound is to recognize, the higher the threshold value may be set. The determination of whether a sound is easy to recognize may be based on, for example, a distribution map of feature information included in each sound identification information in a vector space. In an embodiment, when feature information output corresponding to a specific sound identification information is widely distributed, the sound may be difficult to recognize, while when feature information is densely distributed, the sound may be easy to recognize. That is, a threshold value may be set corresponding to one or more prediction item information. As a specific example, the threshold value for an explosion sound, which is relatively easy to recognize, may be 90, and the threshold value for a child's crying sound, which is difficult to recognize, may be 60. The specific description of the preset threshold values for each sound described above is merely exemplary, and the present invention is not limited thereto. The server 100 can identify one or more recognition sound frames by determining whether each prediction value information is equal to or greater than a predetermined threshold value (S330). Specifically, the server 100 can output a prediction value corresponding to each sound frame. In this case, the prediction value corresponding to each sound frame can include prediction item information and prediction value information.
[0055] As a specific example, the server 100 can identify, through the predicted value corresponding to the first sound frame 411, that the predicted numerical information corresponding to the “sound of breaking glass” and the “sound of a child crying” is “82” and “5”, respectively, and can identify, through the predicted value corresponding to the second sound frame 412, that the predicted numerical information corresponding to the “sound of breaking glass” and the “sound of a siren” is “50” and “12”, respectively.
[0056] In addition, the server 100 may identify a predetermined threshold value for each prediction item information (i.e., the sound of glass breaking, a child crying, and a siren) corresponding to each sound frame. For example, the predetermined threshold values corresponding to the sound of glass breaking, a child crying, and a siren may be identified as 80, 60, and 90, respectively. The server 100 can identify one or more recognition sound frames by comparing the predicted numerical information corresponding to each sound frame with the corresponding threshold value. Specifically, the server 100 can identify one or more recognition sound frames by determining whether each of the predicted numerical information included in the output predicted value is equal to or greater than a predetermined threshold value.
[0057] In this case, the server 100 may identify the first sound frame 411 as one or more recognition sound frames by determining that the predicted numerical information corresponding to the sound of breaking glass in the first sound frame 411 is 82, which is greater than or equal to 80, a predetermined critical value related to the sound of breaking glass. Also, the server 100 may identify the second sound frame 412 as one or more recognition sound frames by determining that the predicted numerical information corresponding to the sound of breaking glass in the second sound frame 412 is 50, which is a predetermined critical value, and is less than 80, a predetermined critical value related to the sound of breaking glass. In other words, the server 100 can identify, as one or more recognized audio frames, only audio frames having prediction numerical information equal to or greater than a predetermined threshold value among one or more audio frames generated based on the audio data 400. That is, the server 100 can remove audio frames associated with less accurate recognition results from each audio frame and identify only frames having a certain reliability or higher as one or more recognized audio frames.
[0058] According to one embodiment of the present invention, the server 100 may identify a converted audio frame through time series analysis based on one or more recognized audio frames (S400). Here, the time series analysis may refer to an analysis that observes the time points at which audio data is acquired and determines whether there is any audio that may be misrecognized. A detailed description of the method for identifying one or more converted audio frames through time series analysis will be provided below with reference to FIG. 7.
[0059] In one embodiment, the server 100 may identify predicted item information corresponding to each of one or more recognition sound frames (S410). That is, the server 100 may identify what kind of sound each of one or more recognition sound frames identified as a result of the threshold value analysis relates to. For example, the one or more recognition sound frames may include a first sound frame, a fourth sound frame, and a fifth sound frame. In this case, the server 100 may identify predicted item information for each sound frame. For example, the predicted item information for the first sound frame may include the "sound of glass breaking," the predicted item information for the fourth sound frame may include the "sound of a siren," and the predicted item information for the fifth sound frame may include the "sound of a siren." The specific descriptions of the one or more recognition sound frames and predicted item information described above are merely examples, and the present invention is not limited thereto. In an embodiment, the server 100 may determine whether the predicted item information is repeated a predetermined number of times or more within a predetermined reference time (S420). Specifically, the server 100 may preset a reference time and a critical number of times for each piece of predicted item information. For example, in the case of the sound of a puppy barking, the reference time may be preset as a time associated with two sound frames, and the critical number of times may be preset to two. In other words, when one or more recognized sound frames are associated with the sound of a puppy barking, the server 100 may identify the predetermined reference time and the predetermined critical number of times for the corresponding item information (i.e., the sound of a puppy barking) and determine whether the sound of a puppy barking has been repeated and recognized a predetermined number of times within the reference time. That is, the server 100 may determine whether a specific sound is continuously recognized a predetermined number of times through one or more recognized sound frames.
[0060] The server 100 can identify a transformed audio frame based on the determination result (S430). If the server 100 determines that a specific audio has not been recognized consecutively for a set reference value through one or more recognition audio frames (i.e., that it has been repeated a predetermined number of times or more within a predetermined reference time), the server 100 can identify at least one of the one or more recognition audio frames as a transformed audio frame. Here, the transformed audio frame may refer to an audio frame that is to be transformed to reduce the probability of misrecognition, i.e., to improve recognition accuracy. According to an embodiment of the present invention, the server 100 may perform a transformation on a predicted value corresponding to the transformed audio frame (S500). In this embodiment, the transformation on the predicted value may include at least one of a noise transformation and an audio item transformation.
[0061] Noise transformation can refer to transforming the output of an acoustic recognition model based on a transformed acoustic frame into an unrecognized item, i.e., transforming the output (i.e., predicted value) of an acoustic recognition model associated with a transformed frame into an unrecognized item (e.g., "others"). The acoustic item transformation may refer to transforming predicted item information associated with the transformed acoustic frame into calibrated predicted item information, where the calibrated predicted item information may be determined based on associations between the predicted item information. As a specific example, if prediction item information related to a converted sound frame includes information about "the sound of washing hands," "the sound of a toilet filling with water," which is associated with the sound of washing hands, may be determined as proofread prediction item information. In this case, the server 100 may convert the prediction item information so that the converted sound frame related to the sound of washing hands is recognized as the sound of a toilet filling with water. The above-described specific descriptions of the prediction item information and proofread prediction item information are merely examples, and the present invention is not limited thereto. As a result, if the server 100 determines that a particular sound has not been recognized consecutively for a set reference number of times through one or more recognized sound frames (i.e., that it has been repeated a predetermined number of times or more within a predetermined reference time), it can identify a conversion frame and perform conversion on the predicted value of the corresponding conversion frame. In this case, converting the predicted value of the conversion frame can mean converting the conversion frame so that it is not recognized (i.e., converting it to an item not targeted for recognition) or converting it so that it is recognized as another sound that does not cause a recognition error when attempting to recognize an event. Such conversion can be possible when one or more sound frames partially overlap with adjacent sound frames. For example, a sound frame that is recognized alone in one sound frame because it is recognized to partially overlap through a second time unit can be identified as a conversion frame and converted. In other words, when trying to detect an event by targeting a specific sound, the recognition accuracy of the voice data can be improved by correcting (or converting) the sound frames related to misrecognition so that they do not cause recognition errors.
[0062] According to one embodiment of the present invention, the server 100 can identify a correlation between each of the one or more recognition audio frames based on prediction item information corresponding to each of the recognition audio frames. For example, the one or more recognition audio frames can include a first audio frame and a second audio frame. The first audio frame can include prediction item information such as "the sound of flushing a toilet," and the second audio frame can include prediction item information such as "the sound of washing hands." The server 100 can identify a correlation between each of the audio frames. For example, the server 100 can identify a correlation such that the acquisition of the second audio frame is predicted after the acquisition of the first audio frame. In an embodiment, the server 100 may determine whether to adjust the threshold value and the threshold number of times corresponding to each of one or more sound frames based on the correlation. The server 100 may adjust the threshold value and the threshold number of times corresponding to the sound frames based on the correlation between the sound frames. That is, the threshold value and the threshold number of times preset for each sound item may be variably adjusted based on the correlation between the sound frames. As a more specific example, server 100 may be configured to detect an event related to the sound of water flushing a toilet. In this case, a first sound frame acquired based on sound data may include predicted item information of "the sound of water flushing a toilet," and a second sound frame may include predicted item information of "the sound of washing hands." For example, the sound related to the second sound frame may also be related to the sound of water flushing and may be similar to the sound event (i.e., the sound of water flushing a toilet) detected or recognized by server 100. Accordingly, server 100 may identify a correlation between sound frames (i.e., a correlation that the acquisition of the sound of washing hands is predicted after the acquisition of the first sound frame) and adjust a threshold value and a critical count corresponding to a predicted item related to the sound of washing hands.
[0063] For example, the server 100 may adjust the threshold value corresponding to the sound prediction item related to the sound of washing hands from the current 80 to 95. Accordingly, the reference value for determining the sound of washing hands in the process of the threshold value analysis may be increased, thereby further improving the accuracy of recognition. In this case, by setting a higher reference value than before, the probability that the sound frame will be recognized as the sound of washing hands may decrease, thereby improving the accuracy of event recognition related to the sound of flushing a toilet. As another example, the server 100 may adjust the threshold number of times for the sound prediction item related to the sound of washing hands after the sound of flushing the toilet from the current 2 to 5. When the sound of washing hands is recognized alone, it is determined to have been recognized successfully even if it is recognized only twice in a row, but it is determined to have been recognized only if it is acquired five times repeatedly after the related sound (i.e., the sound of flushing the toilet). That is, by variably adjusting the threshold value and the critical number of times according to the correlation between sounds, the next sound frame acquired can be processed as an item not yet recognized (e.g., "others"). In other words, when a first sound frame is recognized, the threshold value and the critical number of times associated with a second sound frame (e.g., the sound of washing hands) related to the corresponding first sound frame (e.g., the sound of flushing a toilet) are adjusted to increase the reference value, so that when a second sound frame is subsequently acquired, it can be processed to be recognized as an item not yet recognized. As a result, the recognition accuracy associated with the first sound frame to be detected can be maximized.
[0064] Figure 8 illustrates an exemplary table for explaining the acoustic data correction process according to an embodiment of the present invention. Figure 9 illustrates an exemplary diagram for explaining the acoustic data correction process according to an embodiment of the present invention. 8 may be a table related to predicted values output by an acoustic recognition model corresponding to a case where five acoustic frames are constructed based on acoustic data. As shown in FIG. 8, the five acoustic frames may include a first acoustic frame corresponding to 0 to 1 second, a second acoustic frame corresponding to 0.5 to 1.5 seconds, a third acoustic frame corresponding to 1 to 2 seconds, a fourth acoustic frame corresponding to 1.5 to 2.5 seconds, and a fifth acoustic frame corresponding to 2 to 3 seconds. In this case, the first time unit 400a may be 1 second, and the second time unit 400b may be 0.5 seconds. The start points of adjacent acoustic frames may be configured to differ by the size of the second time unit 400b, which is smaller than the size of the first time unit 400a, so that each acoustic frame may at least partially overlap with its adjacent counterpart.
[0065] In addition, the prediction item information corresponding to each sound frame and the prediction numerical information corresponding to each prediction item information may be as shown in FIG. 8. For example, the closer the prediction numerical information is to 1, the higher the prediction probability, and the closer it is to 0, the lower the prediction probability. For example, it can be confirmed that the output of siren corresponding to 0.5 to 1.5 seconds, i.e., the second sound frame, is 0.9, which is the highest. This may mean that there is a very high probability that the sound acquired between 0.5 and 1.5 seconds is siren. 9(a) is an example diagram showing the results of performing a threshold analysis on the predicted values of FIG. 8. Referring to FIG. 9(a), it can be seen that in the case of "siren," only frames (i.e., the second, third, and fourth sound frames) that are equal to or greater than a threshold value (e.g., 0.6) are identified. In addition, it can be seen that in the case of "scream," only frames (i.e., the first and fourth sound frames) that are equal to or greater than a threshold value (e.g., 0.3) are identified. In addition, it can be seen that in the case of "glass break," only frames (i.e., the fifth sound frame) that are equal to or greater than a threshold value (e.g., 0.7) are identified. For example, sound frames equal to or greater than a threshold value for each predicted item can be identified as one or more recognized sound frames.
[0066] Figure 9(b) is an example diagram showing the results of performing time series analysis corresponding to the predicted values of Figure 8. In this case, the predetermined reference time may be preset to a time associated with two acoustic frames, and the critical number may be preset to two. Referring to (a) and (b) of Figure 9, in the case of siren, it can be seen that when the recognition result of a recognition acoustic frame is observed twice consecutively, only the related frame remains as a recognition target.
[0067] Specifically, in FIG. 9(a), the second sound frame is identified as one or more recognized sound frames, but the time series analysis results show that it has been converted as shown in FIG. 9(b). That is, because the first and second sound frames are not observed twice consecutively in FIG. 9(a), the server 100 can identify the second sound frame as a converted frame and perform a correction to convert it to "others." Accordingly, as shown in FIG. 9(b), an "x" may be displayed in the area of the second sound frame for "siren." This may indicate that the sound of "siren" was not recognized in the corresponding section. Since the second and third sound frames all have prediction numerical information above the threshold value at a later time point, it can be determined that "siren" has been recognized in the third sound frame. Similarly, since the third and fourth sound frames all have prediction numerical information above the threshold value, it can be determined that "siren" has been recognized in the fourth sound frame.
[0068] In addition, in the case of a scream, the threshold analysis result indicates that a scream was detected to have occurred in relation to the first and fourth sound frames, as shown in Figure 9(a). However, since no scream was observed twice consecutively in either the third or fourth sound frame during the time series analysis process, the server 100 may identify the fourth sound frame as a conversion frame and perform a correction to convert it to "others." Accordingly, as shown in Figure 9(b), an "x" may be displayed in the fourth sound frame area of the scream. This may indicate that the siren sound was not recognized in the corresponding section. In a further embodiment, the server 100 may provide information on the recognition result of the entire sound data to the user terminal 200. That is, the information on the recognition result of the entire sound data may include information on what sound was recognized at each time point (e.g., each sound frame) corresponding to the entire sound data acquired in time series. For example, the information on the recognition result of the entire sound data may be the same as (c) of Figure 9.
[0069] Referring to (c) of Figure 9, information indicating that "siren" has been recognized in relation to the second sound frame may be displayed. In this case, the second sound frame related to "siren" may be converted (or corrected) because the sound "siren" was not recognized in the time series analysis process, as shown in (b) of Figure 9. In this embodiment, when providing information on the recognition results of the entire sound data, the server 100 may restore results that exceeded the threshold but were excluded in the time series analysis process. This may be to reflect the corresponding recognition results when providing overall recognition information, since in the sound recognition process, a sound is only used as a recognition target if it is recognized two or more times consecutively.
[0070] The steps of a method or algorithm described in connection with the embodiments of the present invention may be embodied directly in hardware, in a software module executed by hardware, or in a combination thereof. The software module may reside in Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), Flash Memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium commonly known in the art to which the present invention pertains. The components of the present invention may be embodied as a program (or application) stored on a medium for execution in conjunction with a computer, which is hardware. The components of the present invention may be implemented as software programs or software elements. Similarly, embodiments may be implemented in programming or scripting languages such as C, C++, Java, assembler, etc., including various algorithms embodied as a combination of data structures, processes, routines, or other programming constructs. Functional aspects may be embodied as algorithms executed by one or more processors. Although the embodiments of the present invention have been described above with reference to the accompanying drawings, those skilled in the art will understand that the present invention may be embodied in other specific forms without changing the technical spirit or essential characteristics thereof. Therefore, the above-described embodiments should be understood to be illustrative in all respects and not restrictive.
Claims
1. 1. A method performed by one or more processors of a computing device, comprising: constructing one or more acoustic frames based on the acoustic data; processing each of the one or more acoustic frames as an input to an acoustic recognition model and outputting a prediction corresponding to each acoustic frame; identifying one or more recognition sound frames through a threshold value analysis based on the predicted values corresponding to each sound frame; identifying a transformed acoustic frame through time series analysis based on the one or more recognition acoustic frames; and performing a transformation corresponding to the prediction on the transformed acoustic frame; The predicted value is One or more pieces of forecast item information and forecast numerical information corresponding to each of the one or more pieces of forecast item information are included, The critical value analysis and determining whether each of the one or more predicted numerical information corresponding to each of the sound frames is equal to or greater than a predetermined threshold value corresponding to each of the predicted item information, thereby identifying the one or more recognized sound frames; The step of identifying the transformed acoustic frames through time series analysis comprises: identifying predicted item information corresponding to each of the one or more recognition acoustic frames; determining whether the identified predictive item information is repeated a predetermined number of times or more during a predetermined reference time period; and identifying the converted audio frame based on a result of determining whether the identified predicted item information is repeated a predetermined number of times or more during a predetermined reference time; The method further comprises: identifying a correlation between each of the one or more recognition acoustic frames based on predicted item information corresponding to each of the recognition acoustic frames; and determining whether to adjust a critical value and a critical number of times corresponding to each of the one or more audio frames based on the correlation;
2. The step of constructing one or more acoustic frames based on the acoustic data comprises:
2. The method for improving recognition accuracy of acoustic data according to claim 1, further comprising: dividing the acoustic data into predetermined first time unit sizes to form the one or more acoustic frames.
3. The start time of each of the one or more acoustic frames is 3. The method for improving recognition accuracy of acoustic data according to claim 2, wherein the difference between the start time points of adjacent acoustic frames and the magnitude of the second time unit is determined.
4. The transformation for the predicted value is 2. The method for improving the recognition accuracy of acoustic data according to claim 1, further comprising at least one of a noise transformation for converting an output of the acoustic recognition model based on the converted acoustic frame into an unrecognized item, and an acoustic item transformation for converting predicted item information related to the converted acoustic frame into calibrated predicted item information.
5. The calibration prediction item information is The method for improving the recognition accuracy of acoustic data according to claim 4, wherein the prediction item information is determined based on a correlation between the prediction items.
6. memory storing one or more instructions; and a processor that executes the one or more instructions stored in the memory; 10. An apparatus, wherein the processor performs the method of claim 1 by executing the one or more instructions.
7. A computer program stored on a computer-readable recording medium so that the computer program can be combined with a computer that is hardware to perform the method according to claim 1.
Citation Information
Patent Citations
Scream detection device
JP2012048173A
Voice recognition accuracy deterioration factor estimation device, voice recognition accuracy deterioration factor estimation method and program
JP2019139010A
Apparatus for dectecting aucoustic event and method and operating method thereof
KR1020140143069A
Non-audible murmur input alarm device, method, and program
WO2008007616A1