Method, apparatus, and program for providing matching information through audio information analysis
Patent Information
- Application Number
- JP2023570375
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-23
- Filing Date
- 2022-04-07
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2042-04-07
AI Technical Summary
Conventional online advertising methods struggle to adapt to changing user interests and provide targeted advertisements effectively, as they rely on pre-set user information that does not account for real-time changes in user preferences.
An acoustic information analysis method that identifies user characteristics through verbal and non-verbal sounds to generate personalized matching information, including user characteristic information and providing optimized advertisements based on this analysis.
Enhances advertising effectiveness by delivering targeted advertisements that align with users' current interests, reducing costs, and improving user satisfaction by providing relevant information.
Smart Images

Figure 00000025_0000 
Figure 00000026_0000 
Figure 00000027_0000
Abstract
Description
[Technical field]
[0001] The present invention relates to a method for providing matching information suitable for a user, and more particularly, to a technique for providing matching information optimized for a user through analysis of acoustic information. [Background technology]
[0002] Due to the increased use of various electronic devices such as smart TVs, smartphones, tablet PCs, etc. and the activation of Internet services, advertisements provided through electronic devices or online are on the rise.
[0003] For example, as an advertising method using electronic devices or online, there is a method in which the advertising content set by the advertiser on each site is provided in the form of a banner to all users who visit the site. As a specific example, recently, an advertising audience target that receives advertisements online is set, and a custom-made advertisement is provided to the advertising audience target.
[0004] As a method for targeting advertisement viewers to provide customized advertisements, there is a method for identifying the subscriber's areas of interest by acquiring and analyzing subscriber information online.
[0005] This type of online advertising is a restrictive type of advertising, since it collects subscriber information when the subscriber accesses a specific site and uses the service. Also, since online advertising is provided based on user information that is preset as the user's field of interest, if the user's field of interest changes, the user must change the field of interest that he or she has set himself or herself in order to receive advertisements related to the previously set field of interest, and therefore information related to the user's new field of interest cannot be provided.
[0006] Therefore, in the conventional advertising method, it is difficult to improve the efficiency of advertisement provision in terms of providing advertisements according to the subscriber's field of interest, and there is a risk that such advertising method cannot actively respond to the user's field of interest that changes over time. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Republic of Korea Registered Patent 10-2044555 Summary of the Invention [Problem to be solved by the invention]
[0008] The present invention aims to solve the above-mentioned problems by providing a user with more suitable matching information through analysis of acoustic information.
[0009] The problems to be solved by the present invention are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0010] In order to solve the above-mentioned problems, a method for providing matching information through audio information analysis according to various embodiments of the present invention is disclosed, which includes the steps of acquiring audio information, acquiring user characteristic information based on the audio information, and providing matching information corresponding to the user characteristic information.
[0011] In an alternative embodiment, the step of acquiring user characteristic information based on the acoustic information may include a step of identifying a user or object through analysis of the acoustic information, and a step of generating the user characteristic information based on the identified user or object.
[0012] In an alternative embodiment, the step of acquiring user characteristic information based on the acoustic information may include a step of generating activity time information related to a time when a user is active in a specific space through analysis of the acoustic information, and a step of generating the user characteristic information based on the activity time information.
[0013] In an alternative embodiment, the step of providing matching information corresponding to the user characteristic information may include a step of providing the matching information based on the acquisition time and frequency of each of a plurality of user characteristic information corresponding to each of a plurality of acoustic information acquired during a predetermined period of time.
[0014] In an alternative embodiment, the step of acquiring the acoustic information may include a step of performing pre-processing on the acquired acoustic information and a step of identifying acoustic feature information corresponding to the pre-processed acoustic information, and the acoustic feature information may include first feature information regarding whether the acoustic information is related to at least one of linguistic sounds and non-linguistic sounds and second feature information related to object classification.
[0015] In an alternative embodiment, the step of acquiring user characteristic information based on the acoustic information includes a step of acquiring the user characteristic information based on the acoustic characteristic information corresponding to the acoustic information, and the step of acquiring the user characteristic information may include at least one of a step of processing the acoustic information as an input to a first acoustic model to acquire user characteristic information corresponding to the acoustic information when the acoustic characteristic information includes first characteristic information that the acoustic characteristic information is related to linguistic sounds, or a step of processing the acoustic information as an input to a second acoustic model to acquire user characteristic information corresponding to the acoustic information when the acoustic characteristic information includes first characteristic information that the acoustic characteristic information corresponds to non-linguistic sounds.
[0016] In an alternative embodiment, the first acoustic model is a neural network model trained to perform an analysis on acoustic information related to linguistic acoustics to discriminate at least one of text, subject, or emotion associated with the acoustic information, and the second acoustic model is a neural network model trained to perform an analysis on acoustic information related to non-linguistic acoustics to acquire object identification information or object state information associated with the acoustic information, and the user characteristic information may include at least one of first user characteristic information related to at least one of text, subject, or emotion associated with the acoustic information and second user characteristic information related to the object identification information or the object state information associated with the acoustic information.
[0017] In an alternative embodiment, the step of providing matching information corresponding to the user characteristic information may include, when the acquired user characteristic information includes the first user characteristic information and the second user characteristic information, a step of acquiring associated information related to an association between the first user characteristic information and the second user characteristic information, a step of updating matching information based on the associated information, and a step of providing the updated matching information.
[0018] In an alternative embodiment, the step of providing matching information based on the user characteristic information includes the steps of generating an environmental characteristic table based on one or more user characteristic information corresponding to each of one or more pieces of acoustic information acquired at a predetermined time period, and providing the matching information based on the environmental characteristic table, and the environmental characteristic table may be information regarding statistics of each piece of user characteristic information acquired at the predetermined time period.
[0019] In an alternative embodiment, the step of providing matching information corresponding to the user characteristic information may include identifying a first time point for providing the matching information based on the environmental characteristic table, and providing the matching information corresponding to the first time point.
[0020] According to another embodiment of the present invention, there is disclosed an apparatus for performing a method for providing matching information through audio information analysis, the apparatus including a memory for storing one or more instructions and a processor for executing the one or more instructions stored in the memory, the processor executing the one or more instructions to perform the method for providing matching information through audio information analysis.
[0021] According to yet another embodiment of the present invention, there is disclosed a computer program stored on a computer-readable recording medium, which is combined with a computer as hardware to perform the method of providing matching information through audio information analysis described above. Further details of the invention are contained in the detailed description and the drawings. Effect of the Invention
[0022] According to various embodiments of the present invention, it is possible to provide targeted matching information based on audio information acquired in relation to a user's living environment, thereby maximizing advertising effectiveness. The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief description of the drawings]
[0023] [Figure 1] 1 is a diagram illustrating a system for performing a method for providing matching information through audio information analysis according to an embodiment of the present invention. [Diagram 2] 1 is a hardware configuration diagram of a server for providing matching information through audio information analysis in accordance with an embodiment of the present invention. [Diagram 3] 1 illustrates a flowchart showing an example of a method for providing matching information through audio information analysis in accordance with an embodiment of the present invention. [Figure 4]4 is a diagram illustrating an example of a process of acquiring user characteristic information based on acoustic information in accordance with an embodiment of the present invention; [Diagram 5] 1 is a diagram illustrating an example of a process of providing matching information based on user characteristic information related to an embodiment of the present invention. [Figure 6] 1 is an exemplary diagram illustrating a process of acquiring various acoustic information within a space where a user is located and a method of providing matching information corresponding to the acoustic information according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0024] Various embodiments will now be described with reference to the drawings. Various descriptions are presented herein to provide an understanding of the present invention. However, it will be apparent that such embodiments may be practiced without such specific descriptions. As used herein, the terms "component," "module," "system," and the like refer to computer-related entities, hardware, firmware, software, combinations of software and hardware, or software implementations. For example, a component may be, but is not limited to, a procedure running on a processor, a processor, an object, a thread of execution, a program, and / or a computer. For example, both an application running on a computing device and the computing device may be a component. One or more components may reside within a processor and / or thread of execution. A component may be localized within one computer. A component may be distributed among two or more computers. Such components may also execute from a variety of computer-readable media having various data structures stored therein. Components may communicate through local and / or remote processes, for example, by signals comprising one or more data packets (e.g., data from one component interacting with other components in a local system, a distributed system, and / or data transmitted over a network such as the Internet to other systems via signals).
[0025] Additionally, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, it is intended to mean one of the natural inclusive permutations of "X utilizes A or B." That is, if X utilizes A; X utilizes B; or X utilizes both A and B, then "X utilizes A or B" may apply in any of these cases. Additionally, the term "and / or" as used herein should be understood to refer to and include all possible combinations of one or more of the associated listed items. Additionally, the terms "comprise" and / or "comprises" should be understood to mean the presence of the relevant feature and / or component. However, the terms "comprise" and / or "comprises" should be understood not to exclude the presence or addition of one or more other features, components and / or groups thereof. Additionally, in the present specification and claims, the singular should generally be construed to mean "one or more," unless otherwise specified or clear from the context to indicate the singular form.
[0026] Those skilled in the art should further appreciate that the various exemplary logical blocks, configurations, modules, circuits, means, logic, and algorithm steps described in connection with the embodiments disclosed herein may be embodied in electronic hardware, computer software, or any combination of both. To clearly illustrate the interchangeability of hardware and software, the various exemplary components, blocks, configurations, means, logic, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is embodied in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in various ways for each particular application. However, such implementation decisions should not be interpreted as departing from the scope of the present invention.
[0027] The description of the embodiments presented is provided to enable one of ordinary skill in the art to make and practice the present invention. Various modifications to these embodiments will be apparent to those of ordinary skill in the art. The generic principles defined herein may be applied to other embodiments without departing from the scope of the present invention, and the present invention is not limited to the embodiments presented herein. The present invention is to be accorded the broadest scope consistent with the principles and novel features disclosed herein.
[0028] In this specification, the term "computer" refers to any type of hardware device including at least one processor, and may also include software configurations that operate on the hardware device in accordance with the embodiment. For example, the term "computer" refers to any type of device, including, but not limited to, a smartphone, tablet PC, desktop, notebook computer, and user clients and applications that run on each device.
[0029] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Although each step described in this specification is described as being performed by a computer, the subject of each step is not limited to this, and depending on the embodiment, at least some of each step may be performed by different devices.
[0030] Here, the method of providing matching information through audio information analysis according to various embodiments of the present invention can provide optimized matching information to each user based on various audio information acquired in the real life of each of the users. For example, the matching information can be information related to advertisement. That is, providing optimized matching information to a user can mean providing an efficient advertisement that enhances the user's purchasing motivation, i.e., optimized advertisement information. That is, the method of providing matching information through audio information analysis according to the present invention can provide customized advertisement information to the corresponding user by analyzing various audio information acquired in the user's living space. Accordingly, from the viewpoint of an advertiser, since an advertisement can be selectively exposed only to potential customers or target customers who are interested in the advertisement, the advertisement cost can be drastically reduced and the advertisement effect can be maximized. Also, from the viewpoint of a consumer, since an advertisement that satisfies his / her interest and need can be provided, the convenience of information search can be increased.
[0031] 1 is a diagram illustrating a system for performing a method for providing matching information through audio information analysis according to an embodiment of the present invention. As illustrated in FIG. 1, the system for performing a method for providing matching information through audio information analysis according to an embodiment of the present invention may include a server 100 for providing matching information through audio information analysis, a user terminal 200, and an external server 300.
[0032] Here, the system for providing matching information through acoustic information analysis illustrated in FIG. 1 is according to one embodiment, and its components are not limited to the embodiment illustrated in FIG. 1 and may be added, modified or deleted as necessary.
[0033] In one embodiment, the server 100 that provides matching information through audio information analysis can acquire audio information and provide optimal matching information by analyzing the acquired audio information. That is, the server 100 that provides matching information through audio information analysis can acquire various audio information related to the real life of a user, and analyze the acquired audio information through an audio model to grasp information related to the user's tendencies and provide optimal matching information for the user.
[0034] According to the embodiment, the server 100 that provides matching information through audio information analysis may include any server implemented by an API (Application Programming Interface). For example, the user terminal 200 may acquire audio information, perform analysis thereon, and provide the server 100 with an audio information recognition result through the API. Here, the audio information recognition result may mean, for example, a feature related to the audio information analysis. For example, the audio information recognition result may be a spectrogram acquired through a Short-Time Fourier Transform (STFT) on the audio information. The spectrogram is used to visualize and grasp sound or waves, and may be a combination of waveform and spectrum features. The spectrogram may show the difference in amplitude as a difference in print density or display hue according to changes in the time axis and frequency axis.
[0035] As another example, the acoustic information recognition result may include a Mel-Spectrogram obtained through a Mel-Filter Bank for a spectrogram. In general, the human cochlea may have different vibrating parts depending on the frequency of the audio data. The human cochlea has a characteristic of being able to detect frequency changes well in a low frequency band and being unable to detect frequency changes well in a high frequency band. Accordingly, the Mel-Spectrogram may be obtained from the spectrogram by using the Mel-Filter Bank so as to have a recognition ability similar to the characteristics of the human cochlea for the audio data. That is, the Mel-Filter Bank may apply a small number of filter banks to a low frequency band and a wider filter bank to a high frequency band. In other words, the user terminal 200 may apply the Mel-Spectrogram to the spectrogram to recognize the acoustic information similar to the characteristics of the human cochlea. That is, the Mel-Spectrogram may include frequency components reflecting the human hearing characteristics.
[0036] The server 100 can provide optimized matching information to a corresponding user based on the acoustic information recognition result acquired from the user terminal 200. In this case, the server 100 receives only the acoustic information recognition result (e.g., spectrogram or mel spectrogram based on which the user's tendency is predicted) from the user terminal 200, so that the user's privacy issue due to the collection of acoustic information can be resolved.
[0037] In one embodiment, an acoustic model (e.g., an artificial intelligence model) may be comprised of one or more network functions, which may be comprised of a collection of interconnected computational units that may generally be referred to as "nodes." Such "nodes" may also be referred to as "neurons." One or more network functions may be comprised of at least one or more nodes. The nodes (or neurons) that make up one or more network functions may be interconnected by one or more "links."
[0038] In an artificial intelligence model, one or more nodes connected through links can form a relationship of input node and output node relative to each other. The concept of input node and output node is relative, and any node that has an output node relationship with one node can be an input node relationship with other nodes, and vice versa. As mentioned above, the input node to output node relationship can be generated around the link. One or more output nodes can be connected to one input node through links, and vice versa.
[0039] In the relationship between an input node and an output node connected through a link, the value of the output node may be determined based on data input to the input node. Here, the node connecting the input node and the output node may have a weight. The weight may be variable and may be changed by a user or an algorithm so that the artificial intelligence model performs a desired function. For example, when one or more input nodes are connected to one output node through respective links, the output node may determine the output node value based on the value input to the input node connected to the output node and the weight set for the link corresponding to each input node.
[0040] As described above, in an AI model, one or more nodes are connected to each other through one or more links to form input node and output node relationships within the AI model. The characteristics of the AI model can be determined by the number of nodes and links, the correlation between the nodes and links, and the weights assigned to each link within the AI model. For example, if there are two AI models that have the same number of nodes and links but different weights between the links, the two AI models can be recognized as different from each other.
[0041] Some of the nodes constituting the artificial intelligence model may constitute a layer based on the distance from the initial input node. For example, a set of nodes whose distance from the initial input node is n may constitute an n-th layer. The distance from the initial input node may be defined by the minimum number of links that must be traversed to reach the corresponding node from the initial input node. However, such a definition of a layer is arbitrary for the purpose of explanation, and the order of a layer in an artificial intelligence model may be defined in a manner different from that described above. For example, the layer of a node may be defined by the distance from a final output node.
[0042] The first input node may refer to one or more nodes to which data is directly input without passing through a link in relation to other nodes in the artificial intelligence model. Or, in relation between nodes based on a link in the artificial intelligence model network, it may refer to a node that does not have other input nodes connected to a link. Similarly, the final output node may refer to one or more nodes that do not have an output node in relation to other nodes in the artificial intelligence model. Also, the hidden node may refer to a node that constitutes the artificial intelligence model other than the first input node and the final output node. The artificial intelligence model according to an embodiment of the present invention may have more nodes in the input layer than the nodes in the hidden layer close to the output layer, and may be an artificial intelligence model in which the number of nodes decreases as it progresses from the input layer to the hidden layer.
[0043] The artificial intelligence model may include one or more hidden layers. The hidden nodes of the hidden layers may receive the output of the previous layer and the output of the surrounding hidden nodes as input. The number of hidden nodes in each hidden layer may be the same or different. The number of nodes in the input layer may be determined based on the number of data fields of the input data, and may be the same or different from the number of hidden nodes. Input data input to the input layer may be operated by the hidden nodes of the hidden layer and may be output by a fully connected layer (FCL), which is an output layer.
[0044] In various embodiments, the artificial intelligence model may be subjected to supervised learning using a plurality of pieces of acoustic information and specific information corresponding to each piece of acoustic information as learning data, but is not limited thereto and various learning methods may be applied.
[0045] Here, supervised learning is a method of labeling specific data and information related to the specific data to generate learning data and learning using the same, and refers to a method of labeling two pieces of data that have a causal relationship to generate learning data and learning through the generated learning data.
[0046] In one embodiment, the server 100 that provides matching information through acoustic information analysis may determine whether or not to interrupt learning by using validation data when learning of one or more network functions is performed for a predetermined epoch or more. The predetermined epoch may be a part of the entire learning target epoch.
[0047] The verification data may be composed of at least a part of the labeled learning data. That is, the server 100 that provides matching information through acoustic information analysis performs learning of an artificial intelligence model through the learning data, and after the learning of the artificial intelligence model is repeated for a predetermined epoch or more, the server 100 that provides matching information through acoustic information analysis may determine whether the learning effect of the artificial intelligence model is equal to or higher than a predetermined level using the verification data. For example, when performing learning with a target number of repeated learning times of 10 using 100 pieces of learning data, the server 100 that provides matching information through acoustic information analysis may perform 3 repeated learning times using 10 pieces of verification data after performing 10 repeated learning times, which are the predetermined epochs, and if the change in the output of the artificial intelligence model during the three repeated learning times is equal to or lower than a predetermined level, the server 100 may determine that further learning is meaningless and terminate the learning.
[0048] That is, the validation data may be used to determine the completion of learning based on whether the effect of the learning per epoch is above or below a certain level in the iterative learning of the artificial intelligence model. The numbers of learning data, validation data, and the number of iterations described above are merely examples, and the present invention is not limited thereto.
[0049] The server 100 that provides matching information through acoustic information analysis may generate an AI model by testing the performance of one or more network functions using test data and determining whether or not to activate one or more network functions. The test data may be used to verify the performance of the AI model and may be composed of at least a portion of the learning data. For example, 70% of the learning data may be used for learning the AI model (i.e., learning to adjust weights to output result values similar to the labels), and 30% may be used as test data to verify the performance of the AI model. The server 100 that provides matching information through acoustic information analysis may input the test data to the AI model that has completed learning, measure an error, and determine whether or not to activate the AI model depending on whether the error is greater than or equal to a predetermined performance.
[0050] The server 100, which provides matching information through acoustic information analysis, verifies the performance of the AI model that has completed training using test data for the AI model, and if the performance of the AI model that has completed training is above a predetermined standard, it can activate the AI model for use in other applications.
[0051] In addition, the server 100 that provides matching information through acoustic information analysis can deactivate and discard the AI model if the performance of the AI model that has completed learning is below a predetermined standard. For example, the optimal stimulation position calculation server 100 can determine the performance of the generated AI model based on factors such as accuracy, precision, and recall. The above-mentioned performance evaluation criteria are merely examples, and the present invention is not limited thereto. According to an embodiment of the present invention, the optimal stimulation position calculation server 100 can generate a plurality of AI models by learning each AI model independently, and can evaluate the performance and use only AI models that have a certain level of performance or higher. However, the present invention is not limited thereto.
[0052] Throughout this specification, the terms computational model, neural network, network function, and neural network may be used interchangeably (hereinafter, the term neural network will be used interchangeably). The data structure may include a neural network. The data structure including the neural network may be stored on a computer-readable medium. The data structure including the neural network may also include data input to the neural network, weights of the neural network, hyperparameters of the neural network, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for training the neural network. The data structure including the neural network may include any of the components disclosed above. That is, the data structure including the neural network may include all or any combination of data input to the neural network, weights of the neural network, hyperparameters of the neural network, data obtained from the neural network, activation functions associated with each node or layer of the neural network, and loss functions for training the neural network. In addition to the above-mentioned configurations, the data structure including the neural network may include any other information that determines the characteristics of the neural network. Additionally, the data structure may include any type of data used or generated in the computational process of the neural network, and is not limited to the above. The computer readable medium may include a computer readable recording medium and / or a computer readable transmission medium. A neural network may be comprised of a collection of interconnected computational units that may generally be referred to as nodes. Such nodes may be referred to as neurons. A neural network is comprised of at least one or more nodes.
[0053] According to an embodiment of the present invention, the server 100 that provides matching information through acoustic information analysis may be a server that provides a cloud computing service. More specifically, the server 100 that provides matching information through acoustic information analysis may be a server that provides a cloud computing service that is a type of Internet-based computing and processes information not on the user's computer but on another computer connected to the Internet. The cloud computing service may be a service that stores data on the Internet and allows users to use the data or programs they need through an Internet connection anytime and anywhere without installing them on their own computers, and allows users to easily share and transfer data stored on the Internet with simple operations and clicks. In addition, the cloud computing service may be a service that not only simply stores data on a server on the Internet, but also allows users to perform desired tasks using the functions of applications provided on the web without installing separate programs, and allows multiple people to work while sharing documents at the same time. In addition, the cloud computing service may be embodied in at least one form of IaaS (Infrastructure as a Service), PaaS (Platform as a Service), SaaS (Software as a Service), a virtual machine-based cloud server, and a container-based cloud server. That is, the server 100 that provides matching information through audio information analysis of the present invention may be embodied in the form of at least one of the above-mentioned cloud computing services. The above-mentioned specific descriptions of the cloud computing services are merely examples, and the present invention may include any platform that constructs a cloud computing environment.
[0054] In various embodiments, the server 100 that provides matching information through acoustic information analysis may be connected to the user terminal 200 through a network 400, and may not only generate and provide an acoustic model that analyzes acoustic information, but also provide optimal matching information corresponding to each user based on information obtained by analyzing the acoustic information through the acoustic model (e.g., user characteristic information).
[0055] Here, the network 400 may refer to a connection structure that allows information exchange between each of a plurality of nodes such as terminals and servers, etc. For example, the network 400 may include a local area network (LAN), a wide area network (WAN), the Internet (WWW), wired / wireless data communication networks, telephone networks, wired / wireless television communication networks, etc.
[0056] In addition, the wireless data communication network includes, but is not limited to, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), 5GPP (5th Generation Partnership Project), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, the Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), RF (Radio Frequency), Bluetooth network, NFC (Near-Field Communication) network, satellite broadcasting network, analog broadcasting network, DMB (Digital Multimedia Broadcasting) network, etc.
[0057] In one embodiment, the user terminal 200 may be connected to the server 100 that provides matching information through acoustic information analysis through the network 400, and may provide a plurality of acoustic information (e.g., linguistic acoustic information or non-linguistic acoustic information) to the server 100 that provides matching information through acoustic information analysis, and may receive various information (e.g., user characteristic information corresponding to the acoustic information and matching information corresponding to the user characteristic information, etc.) in response to the provided acoustic information.
[0058] Here, the user terminal 200 is a wireless communication device that is guaranteed to be portable and mobile, and may include, but is not limited to, all kinds of handheld-based wireless communication devices such as navigation, personal communication system (PCS), global system for mobile communications (GSM), personal digital cellular (PDC), personal handyphone system (PHS), personal digital assistant (PDA), international mobile telecommunication (IMT)-2000, code division multiple access (CDMA)-2000, W-code division multiple access (W-CDMA), wireless broadband internet (Wibro) terminal, smartphone, smartpad, tablet PC, etc. For example, the user terminal 200 may further include an artificial intelligence (AI) speaker and an artificial intelligence TV that provide various functions such as music listening and information search through interaction with a user based on a hot word.
[0059] In one embodiment, the user terminal 200 may include a first user terminal 210 and a second user terminal 220. The user terminals (first user terminal 210 and second user terminal 220) may have a mechanism for communicating with each other or with other entities through the network 400 and may refer to any type of entity in a system for providing matching information through audio information analysis. As an example, the first user terminal 210 may include any terminal associated with a user who receives matching information. Also, the second user terminal 220 may include any terminal associated with an advertiser for registering matching information. Such a user terminal 200 may have a display, receive a user's input, and provide any type of output to the user.
[0060] In one embodiment, the external server 300 may be connected to the server 100 that provides matching information through acoustic information analysis via a network 400, and may provide various information / data necessary for the server 100 that provides matching information through acoustic information analysis to analyze acoustic information using an artificial intelligence model, or may receive, store, and manage result data derived by performing acoustic information analysis using an artificial intelligence model. For example, the external server 300 may be, but is not limited to, a storage server separately provided outside the server 100 that provides matching information through acoustic information analysis. Hereinafter, a hardware configuration of the server 100 that provides matching information through acoustic information analysis will be described with reference to FIG. 2.
[0061] FIG. 2 is a hardware configuration diagram of a server for providing matching information through audio information analysis in accordance with an embodiment of the present invention. 2, an optimal stimulation position calculation server 100 (hereinafter, "server 100") according to an embodiment of the present invention may include one or more processors 110, a memory 120 for loading a computer program 151 executed by the processor 110, a bus 130, a communication interface 140, and a storage 150 for storing the computer program 151. Here, only components related to the embodiment of the present invention are illustrated in FIG. 2. Therefore, a person skilled in the art to which the present invention pertains will understand that other general components may be included in addition to the components illustrated in FIG. 2.
[0062] The processor 110 controls the overall operation of each component of the server 100. The processor 110 may include a central processing unit (CPU), a microprocessor unit (MPU), a microcontroller unit (MCU), a graphic processing unit (GPU), or any other type of processor commonly known in the technical field of the present invention.
[0063] The processor 110 can read a computer program stored in the memory 120 to perform data processing for an artificial intelligence model according to an embodiment of the present invention. According to an embodiment of the present invention, the processor 110 can perform calculations for learning a neural network. The processor 110 can perform calculations for learning a neural network, such as processing input data for learning using deep learning (DL), extracting features from the input data, calculating errors, and updating weights of the neural network using backpropagation.
[0064] In addition, at least one of the processor 110, a CPU, a GPGPU, and a TPU, may process the learning of a network function. For example, the CPU and the GPGPU may both process the learning of a network function and the data classification using the network function. In addition, in one embodiment of the present invention, the processors of a plurality of computing devices may be used together to process the learning of a network function and the data classification using the network function. In addition, the computer program executed in the computing device according to one embodiment of the present invention may be a CPU, GPGPU, or TPU executable program.
[0065] As used herein, a network function may be used interchangeably with an artificial neural network or a neural network. As used herein, a network function may include one or more neural networks, in which case the output of the network function may be an ensemble of the outputs of the one or more neural networks.
[0066] The processor 110 may read a computer program stored in the memory 120 to provide an acoustic model according to an embodiment of the present invention. According to an embodiment of the present invention, the processor 110 may obtain user characteristic information corresponding to the acoustic information. According to an embodiment of the present invention, the processor 110 may perform calculations for training the acoustic model.
[0067] According to one embodiment of the present invention, the processor 110 may generally process the overall operation of the server 100. The processor 110 may process signals, data, information, etc. input or output through the components detailed above, or may run application programs stored in the memory 120, thereby providing or processing appropriate information or functions to a user or a user terminal.
[0068] Furthermore, the processor 110 may execute operations for at least one application or program for executing a method according to an embodiment of the present invention, and the server 100 may include one or more processors.
[0069] In various embodiments, the processor 110 may further include a random access memory (RAM, not shown) and a read-only memory (ROM, not shown) for temporarily and / or permanently storing signals (or data) processed within the processor 110. The processor 110 may also be implemented in the form of a system on chip (SoC) including at least one of a graphics processing unit, a RAM, and a ROM.
[0070] The memory 120 stores various data, instructions and / or information. The memory 120 can load a computer program 151 from the storage 150 to perform the methods / operations according to various embodiments of the present invention. When the computer program 151 is loaded into the memory 120, the processor 110 can execute one or more instructions constituting the computer program 151 to perform the methods / operations. The memory 120 may be embodied as a volatile memory such as a RAM, although the scope of the present invention is not limited in this respect.
[0071] The bus 130 provides a communication function between the components of the server 100. The bus 130 may be implemented as various types of buses such as an address bus, a data bus, and a control bus.
[0072] The communication interface 140 supports wired / wireless Internet communication of the server 100. The communication interface 140 may also support various communication methods other than Internet communication. To this end, the communication interface 140 may be configured to include a communication module that is widely known in the technical field of the present invention. In some embodiments, the communication interface 140 may be omitted.
[0073] The storage 150 may non-temporarily store a computer program 151. When the process of providing matching information through audio information analysis is performed through the server 100, the storage 150 may store various information required to provide the process of providing matching information through audio information analysis.
[0074] Storage 150 may be configured to include non-volatile memory such as ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), flash memory, etc., a hard disk, a removable disk, or any form of computer-readable recording medium commonly known in the technical field to which the present invention belongs.
[0075] The computer program 151, when loaded into the memory 120, may include one or more instructions that cause the processor 110 to perform the methods / operations according to various embodiments of the present invention. That is, the processor 110 may execute the one or more instructions to perform the methods / operations according to various embodiments of the present invention.
[0076] In one embodiment, the computer program 151 may include one or more instructions to perform a method of providing matching information through audio information analysis, including the steps of acquiring audio information, acquiring user characteristic information based on the audio information, and providing matching information corresponding to the user characteristic information.
[0077] The steps of a method or algorithm described in connection with the embodiments of the present invention may be embodied directly in hardware, in a software module executed by hardware, or in a combination thereof. The software module may reside in a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable ROM (EPROM), an Electrically Erasable Programmable ROM (EEPROM), a Flash Memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer readable recording medium commonly known in the art to which the present invention pertains.
[0078] The components of the present invention may be embodied as a program (or application) and stored on a medium to be executed in combination with a computer, which is hardware. The components of the present invention may be implemented as software programming or software elements, and similarly, embodiments include various algorithms embodied as a combination of data structures, processes, routines or other programming constructs, and may be embodied in programming or scripting languages such as C, C++, Java, assembler, etc. Functional aspects may be embodied as algorithms executed by one or more processors. Hereinafter, a method for providing matching information through audio information analysis performed by the server 100 will be described with reference to Figures 3 to 6.
[0079] FIG. 3 is a flow chart illustrating an example of a method for providing matching information through audio information analysis according to an embodiment of the present invention.
[0080] According to an embodiment of the present invention, in step S110, the server 100 may perform a step of acquiring acoustic information. According to the embodiment, the acoustic information may be acquired through a user terminal 200 associated with a user. For example, the user terminal 200 associated with a user may include any kind of handheld-based wireless communication device such as a smartphone, a smartpad, a tablet PC, etc., or an electronic device (e.g., a device capable of receiving acoustic information through a microphone) provided in a specific space (e.g., a user's residential space), etc.
[0081] According to an embodiment, the acquisition of acoustic information may involve receiving or loading acoustic information stored in a memory, or may involve receiving or loading acoustic information from another storage medium, another server, or a separate processing module within the same server based on wired / wireless communication means.
[0082] According to a further embodiment of the present invention, the acquisition of the acoustic information may be performed based on whether the user is located in a specific space (e.g., the user's activity space). Specifically, a sensor module may be provided in the specific space related to the user's activity. That is, whether the user is located in the specific space may be identified through the sensor module provided in the specific space. For example, the unique information of the user at a long distance may be recognized using radio waves by using RFID (Radio Frequency Identification) technology, which is a type of short-range communication technology. For example, the user may possess a card or a mobile terminal including an RFID module. Information identifying the user (e.g., the user's personal ID, identification code, etc. registered in the service management server) may be recorded in the RFID module possessed by the user. The sensor module may identify whether the user is located in the specific space by identifying the RFID module possessed by the user. The sensor module may include various technologies (e.g., short-range communication technologies such as Bluetooth) capable of transmitting and receiving unique information of the user in a contact / non-contact manner in addition to RFID technology. In addition, the sensor module may further include a biometric data identification module that identifies the user's biometric data (voice, fingerprint, face) in conjunction with a microphone, touchpad, camera module, etc. In a further embodiment, it may be possible to identify whether the user is located in a particular space through audio information related to the user's speech. Specifically, audio information related to the user's speech may be recognized as a starting word, and additional audio information generated in the space corresponding to the recognition time may be obtained.
[0083] The server 100 can identify whether the user is located in a specific space through the sensor module or the sound associated with the user's speech as described above. In addition, when the server 100 determines that the user is located in a specific space, it can obtain sound information generated at a corresponding time. In other words, if the user is not present in a particular space, audio information related to the space is not acquired, and audio information related to the space is acquired only when the user is present in the particular space, thereby minimizing power consumption.
[0084] According to one embodiment of the present invention, the step of acquiring acoustic information may include a step of performing pre-processing on the acquired acoustic information and a step of identifying acoustic characteristic information corresponding to the pre-processed acoustic information.
[0085] According to an embodiment, the pre-processing of the acoustic information may be a pre-processing for improving the recognition rate of the acoustic information. For example, such pre-processing may include a pre-processing for removing noise from the acoustic information. Specifically, the server 100 may standardize the magnitude of a signal included in the acoustic information based on a comparison between the magnitude of the signal included in the acoustic information and the magnitude of a reference signal. The server 100 may perform a pre-processing for adjusting the magnitude of the signal to be larger when the magnitude of the signal included in the acquired acoustic information is smaller than a predetermined reference signal, and adjusting the magnitude of the signal to be smaller (i.e., not clipped) when the signal included in the acoustic information is equal to or larger than the predetermined reference signal. The above-mentioned specific description of noise removal is merely an example, and the present invention is not limited thereto.
[0086] According to another embodiment of the present invention, the pre-processing of the audio information may include a pre-processing for amplifying sounds other than speech (i.e., non-linguistic sounds) by analyzing the waveform of a signal included in the audio information. Specifically, the server 100 may analyze various sound frequencies included in the audio information and amplify sounds related to at least one specific frequency.
[0087] For example, the server 100 may classify various types of sounds included in the sound information using a machine learning algorithm such as SVM (Supporting Vector Machine) to identify the types of sounds, and amplify specific sounds using a sound amplification algorithm corresponding to each sound including different frequencies. The above-mentioned sound amplification algorithms are merely examples, and the present invention is not limited thereto.
[0088] In other words, the present invention can perform pre-processing to amplify non-verbal sounds included in the acoustic information. For example, in the present invention, the acoustic information used as the basis for analysis to grasp the characteristics of a user (or to provide optimal matching information to a user) can include linguistic acoustic information and non-verbal acoustic information. In one embodiment, non-verbal acoustic information can provide a more meaningful analysis for analyzing the characteristics of a user than linguistic acoustic information.
[0089] As a specific example, when the server 100 acquires acoustic information including the sound (i.e., non-verbal acoustic information) of a companion animal (e.g., a puppy), the server 100 can amplify the acoustic information related to the non-verbal sound in order to improve the recognition rate of the non-verbal acoustic information, “the voice of a puppy.”
[0090] For another example, when the server 100 acquires acoustic information related to a user's cough (i.e., non-verbal acoustic information), the server 100 may amplify the acoustic information related to the non-verbal acoustic information in order to improve the recognition rate of the non-verbal acoustic information, "a person's cough." In other words, by performing pre-processing to amplify non-verbal acoustic information that provides meaningful information for understanding the characteristics of the user, more suitable matching information can be provided to the user.
[0091] The server 100 may also identify acoustic feature information corresponding to the pre-processed acoustic information, where the acoustic feature information may include first feature information regarding whether the acoustic information is related to at least one of linguistic and non-linguistic acoustics, and second feature information related to an object classification.
[0092] The first characteristic information may include information regarding whether the acoustic information is linguistic or non-linguistic. For example, the first characteristic information corresponding to the first acoustic information may include information that the first acoustic information is related to linguistic acoustic, and the first characteristic information corresponding to the second acoustic information may include information that the second acoustic information is related to non-linguistic acoustic.
[0093] The second characteristic information may include information regarding how many objects the audio information includes. For example, the second characteristic information corresponding to the first audio information may include information that the first audio information includes speeches of three users, and the second characteristic information corresponding to the second audio information may include information that the second audio information includes a sound related to the operation of a washing machine and a sound related to a cat's meow. In one embodiment, the first characteristic information and the second characteristic information corresponding to the audio information may be identified through a first audio model and a second audio model, which will be described later.
[0094] That is, the server 100 can perform preprocessing on the acquired acoustic information and identify acoustic characteristic information corresponding to the preprocessed acoustic information. As described above, the acoustic characteristic information includes information on whether the corresponding acoustic information is related to linguistic acoustics or non-linguistic acoustics (i.e., first characteristic information) and information on how many objects exist in the acoustic information (i.e., second characteristic information), which can provide convenience in an acoustic information analysis process described below.
[0095] According to an embodiment of the present invention, in step S120, the server 100 may perform a step of acquiring user characteristic information based on the acoustic information. In one embodiment, the step of acquiring the user characteristic information may include a step of identifying a user or an object through an analysis of the acoustic information, and a step of generating user characteristic information based on the identified user or object.
[0096] Specifically, when the server 100 acquires audio information, it can identify a user or object corresponding to the audio information through an analysis of the audio information. For example, the server 100 can perform an analysis of first audio information and identify that the first audio information is audio corresponding to a first user. As another example, the server 100 can perform an analysis of second audio information and identify that the second audio information is audio corresponding to a vacuum cleaner. As yet another example, the server 100 can perform an analysis of third audio information and identify that the third audio information includes audio corresponding to a second user and audio related to a washing machine. The above-mentioned specific descriptions of the first to third audio information are merely examples, and the present invention is not limited thereto.
[0097] In addition, the server 100 may generate user characteristic information based on a user or object identified in response to the audio information. Here, the user characteristic information is information based on which matching information is provided, and may be information related to text, subject, emotion, object identification information, object state information, etc., related to the audio information.
[0098] For example, when the first sound information is identified as sound corresponding to the first user, the server 100 may identify user information matched to the first user and generate user characteristic information that the user is related to a 26-year-old woman. As another example, when the second sound information is identified as sound corresponding to an A brand vacuum cleaner, the server 100 may generate user characteristic information that the user uses an A brand vacuum cleaner through the second sound information. As yet another example, when the third sound information is identified as sound corresponding to a second user and sound related to a B brand washing machine, the server 100 may generate user characteristic information that the user is related to a 40-year-old man and uses a B brand washing machine. The above-mentioned first to third sound information and the specific description of the user characteristic information corresponding to each sound information are merely examples, and the present invention is not limited thereto. That is, the server 100 may generate user characteristic information related to a user based on a user or object identified through the analysis of the audio information. Such user characteristic information may be information that can grasp the user's interests, preferences, characteristics, etc.
[0099] According to an embodiment, the step of acquiring the user characteristic information may include generating activity time information related to the time when the user is active in a specific space through an analysis of the audio information, and generating the user characteristic information based on the activity time information. Specifically, the server 100 may generate activity time information related to the time when the user is active in the specific space through an analysis of the audio information acquired corresponding to the specific space.
[0100] In one embodiment, the server 100 can obtain information on whether the user is located in a specific space through audio information related to the user's speech. For example, the server 100 can recognize a voice related to the user's speech as a start word, and determine that the user has entered a specific space based on the time of recognition. If audio information acquired in the space does not include a voice related to the user's speech and the volume of the acquired audio information is equal to or less than a predetermined reference value, the server 100 can determine that the user is not in the specific space. The server 100 can also generate activity time information related to the time when the user is active in the specific space based on each determination time. That is, the server 100 can identify whether the user is located in a specific space through audio information related to the user's speech and generate activity time information related to the user.
[0101] In another embodiment, the server 100 may acquire information regarding whether the user is located in a specific space based on the magnitude of the acquired sound information. For example, the server 100 may identify a time point when the magnitude of sound information continuously acquired in the specific space is equal to or greater than a predetermined reference value to determine that the user has entered the specific space, and may identify a time point when the magnitude of sound information acquired in the space is less than a predetermined reference value to determine that the user is not in the specific space. The server 100 may also generate activity time information related to the time when the user is active in the specific space based on each determination time point. That is, the server 100 may identify the magnitude of sound information generated in the specific space, and may identify whether the user is located in the specific space based on the magnitude to generate activity time information related to the user.
[0102] In still another embodiment, the server 100 may obtain information regarding whether the user is located in a specific space based on a specific start sound. Here, the specific start sound may be a sound related to the entry and exit of the user. For example, the start sound may be a sound related to the opening and closing sound of the front door. That is, the server 100 may determine that the user is located in the corresponding space based on sound information related to the opening of the door lock using a password from the outside. Also, the server 100 may determine that the user is not present in the corresponding space based on sound information related to the opening of the front door from the inside. Also, the server 100 may generate activity time information related to the time when the user is active in the specific space based on each determination time point. That is, the server 100 may identify whether the user is located in the specific space based on the start sound generated in the specific space, and generate activity time information related to the user.
[0103] As described above, the server 100 may generate activity time information related to the time when a user is active in a specific space according to various embodiments. For example, the server 100 may generate activity time information indicating that the time when a first user is active in a specific space (e.g., a residential space) corresponds to 18 hours out of 24 hours (e.g., the server 100 is located in the specific space from 12:00 a.m. to 6:00 p.m.). As another example, the server 100 may generate activity time information indicating that the time when a second user is active in a specific space corresponds to 6 hours in a day (e.g., the server 100 is located in the specific space from 12:00 a.m. to 6:00 a.m.). The above-mentioned specific descriptions of the activity time information corresponding to each user are merely examples, and the present invention is not limited thereto.
[0104] Also, the server 100 may generate user characteristic information based on the activity time information. For example, the server 100 may generate user characteristic information that the first user is mostly active in a particular space (e.g., residential space) based on the activity time information that the time the first user is active in the particular space corresponds to 18 hours out of 24 hours. As an example, the server 100 may generate user characteristic information that the first user is related to a "housewife" or a "telecommuter" by identifying that the first user spends more time in the residential space through the activity time information of the first user. In a further embodiment, the server 100 may infer a more specific occupation of the user by combining the activity time information and analysis information on the audio information. In this case, the user's characteristics may be more specifically identified, and the accuracy of the provided matching information may be improved.
[0105] As another example, the server 100 may generate user characteristic information that the second user is less active in the residential space based on activity time information that the time the second user is active in a specific space corresponds to 6 hours per day. The specific description of the activity time information of each user and the user characteristic information corresponding to each activity time information is merely an example, and the present invention is not limited thereto.
[0106] According to another embodiment of the present invention, the step of acquiring user characteristic information may include acquiring user characteristic information based on acoustic characteristic information corresponding to the acoustic information, wherein the acoustic characteristic information may include first characteristic information regarding whether the acoustic information is related to at least one of linguistic acoustics and non-linguistic acoustics, and second characteristic information related to object classification.
[0107] Specifically, the step of acquiring user characteristic information may include at least one of the steps of: if the acoustic characteristic information includes first characteristic information related to linguistic sounds, processing the acoustic information as an input to a first acoustic model to acquire user characteristic information corresponding to the acoustic information; or, if the acoustic characteristic information includes first characteristic information related to non-linguistic sounds, processing the acoustic information as an input to a second acoustic model to acquire user characteristic information corresponding to the acoustic information.
[0108] According to one embodiment, the first acoustic model may be a neural network model trained to perform analysis on acoustic information related to linguistic sounds and to identify at least one of text, subject, or emotion associated with the acoustic information.
[0109] According to an embodiment of the present invention, the first acoustic model is a voice recognition model that receives voice information (i.e., linguistic sound) related to a user's speech included in the acoustic information as an input and outputs text information corresponding to the voice information, and may include one or more network functions pre-trained through training data. That is, the first acoustic model may include a voice recognition model that converts voice information related to a user's speech into text information. For example, the voice recognition model may process voice information related to a user's speech as an input and output corresponding text (e.g., "The puppy's food is out"). The above-mentioned specific descriptions of the voice information and the text corresponding to the voice information are merely examples, and the present invention is not limited thereto.
[0110] In addition, the first acoustic model may include a text analysis model that grasps the subject or emotion contained in the voice information by grasping the context through natural language processing analysis of the text information output corresponding to the voice information.
[0111] In one embodiment, the text analysis model can recognize important keywords and grasp the subject of text information through semantic analysis of the text through a natural language processing neural network (i.e., a text analysis model). For example, in the case of text information related to "puppy food is out," the text analysis model can grasp the subject of the corresponding sentence as "food exhaustion." The above description of the text information and the corresponding subject is merely an example, and the present invention is not limited thereto.
[0112] In addition, in one embodiment, the text analysis model may calculate the text information through a natural language processing neural network and output an analysis value according to each of various intention groups. The various intention groups may be obtained by dividing a sentence including text into specific intentions according to a predetermined criterion. Here, the natural language processing artificial neural network may input the text information as input data, calculate the connection weights of each, and output the intention groups as output nodes. Here, the connection weights may be weights of input, output, and forget gates used in the LSTM method, or weights of gates used in the RNN. Accordingly, the first acoustic model may calculate an analysis value of the text information that corresponds one-to-one to each intention group. Here, the analysis value may mean a probability that the text information can belong to one intention group. In a further embodiment, the first acoustic model may further include a sentiment analysis model that outputs an analysis value for emotions through a voice analysis according to a change in voice pitch. That is, the first acoustic model may include a voice recognition model that outputs text information corresponding to the voice information of a user, a text analysis model that grasps the subject of a sentence through a natural language processing analysis of the text information, and a sentiment analysis model that grasps the emotion of a user through a voice analysis according to a change in voice pitch. Accordingly, the first acoustic model may output information on a text, a subject, or an emotion related to the acoustic information based on the acoustic information including voice information related to the user's speech.
[0113] Also, according to an embodiment, the second acoustic model may be a neural network model trained to perform analysis on acoustic information related to non-linguistic sounds to obtain object identification information and object state information related to the acoustic information.
[0114] According to an embodiment of the present invention, the second acoustic model is a dimension reduction network function and a dimension restoration network function trained by the server 100 to output output data similar to input data, and may be implemented through a trained dimension reduction network function. That is, the second acoustic model may be configured through a dimension reduction network function in the configuration of a trained autoencoder.
[0115] According to an embodiment, the server 100 can train the autoencoder through an unsupervised learning method. Specifically, the server 100 can train a dimension reduction network function (e.g., an encoder) and a dimension restoration network function (e.g., a decoder) that configure the autoencoder to output output data similar to the input data. In more detail, only the core feature data (or features) of the audio information input in the encoding process through the dimension reduction network function can be learned through a hidden layer, and the remaining information can be lost. In this case, the output data of the hidden layer in the decoding process through the dimension restoration network function can be an approximation of the input data (i.e., audio information) rather than a perfect copy value. That is, the server 100 can train the autoencoder by adjusting the weights so that the output data and the input data are as similar as possible.
[0116] An autoencoder may be a type of neural network for outputting output data similar to input data. An autoencoder may include at least one hidden layer, and an odd number of hidden layers may be disposed between the input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called a bottleneck layer (encoding), and then expanded symmetrically from the bottleneck layer to the output layer (symmetric to the input layer). The number of input layers and output layers may correspond to the number of items of input data remaining after preprocessing of the input data. In the autoencoder structure, the number of nodes in the hidden layer included in the encoder may decrease as it moves away from the input layer. The number of nodes in the bottleneck layer (the layer with the fewest nodes located between the encoder and the decoder) may be maintained at a certain number or more (e.g., more than half of the input layer) because if the number is too small, a sufficient amount of information may not be transmitted.
[0117] The server 100 may use a training data set including a plurality of training data tagged with object information as an input of a trained dimension reduction network to match the output feature data for each object with the tagged object information and store the matched feature data. Specifically, the server 100 may use a dimension reduction network function to obtain feature data of a first object for the training data included in the first training data subset by using a first training data subset tagged with first object information (e.g., a puppy) as an input of the dimension reduction network function. The obtained feature data may be expressed as a vector. In this case, the feature data output corresponding to each of the plurality of training data included in the first training data subset is an output through the training data related to the first object, so that they may be located relatively close to each other in the vector space. The server 100 may match the first object information (i.e., a puppy) with the feature data related to the first object expressed as a vector and store the matched feature data. In the case of a dimension reduction network function of a trained autoencoder, a dimension restoration network function may be trained to extract features that allow the input data to be restored well. Therefore, the second acoustic model is implemented through a dimensionality reduction network function in the learning autoencoder, thereby extracting features (i.e., the acoustic style of each object) that can well restore the input data (e.g., acoustic information).
[0118] For a further example, a plurality of training data included in each of the second training data subsets tagged with second object information (e.g., cat) may be converted into feature data through a dimensionality reduction network function and displayed in a vector space. In this case, the feature data may be output through the training data related to the second object information (i.e., cat), and therefore may be located relatively close to each other in the vector space. In this case, the feature data corresponding to the second object information may be displayed in a vector space different from the feature data corresponding to the first object information.
[0119] That is, when acoustic information generated in a specific space (e.g., a living space) is input, the dimension reduction network function constituting the second acoustic model through the above-mentioned learning process can extract features corresponding to the acoustic information by calculating the corresponding acoustic information using the dimension reduction network function. In this case, the second acoustic model can evaluate the similarity of the acoustic style by comparing the distance in the vector space between the area in which the feature corresponding to the acoustic information is displayed and the object-specific feature data, and can obtain object identification information or object state information corresponding to the acoustic information based on the corresponding similarity evaluation.
[0120] Specifically, the second acoustic model may acquire the first feature information by calculating the first acoustic information received from the first user terminal using a dimension reduction network function. In this case, the second acoustic model may acquire object identification information or object state information corresponding to the first acoustic information based on a position between the first feature information and object-specific feature data pre-recorded on a vector space through learning.
[0121] As a specific example, object identification information that the first sound information is related to "A brand washing machine" can be acquired based on the space of the first object (e.g., A brand washing machine) having the closest distance in the vector space to the first feature information. In another example, object state information that the first sound information is related to "a person's coughing sound" may be obtained based on a second object (e.g., a person's coughing sound) space having a distance in the vector space closest to the first feature information. The above-mentioned specific descriptions of the object identification information and object state information are merely examples, and the present invention is not limited thereto.
[0122] 4, when acquiring acoustic information, the server 100 may determine whether the corresponding acoustic information is related to linguistic acoustics or non-linguistic acoustics based on the first characteristic information. When the acoustic information corresponds to linguistic acoustics, the server 100 may determine text, subject, or emotion corresponding to the acoustic information using the first acoustic model, and acquire first user characteristic information based on the corresponding information. When the acoustic information corresponds to non-linguistic acoustics, the server 100 may acquire object identification information or object state information corresponding to the acoustic information using the second acoustic model, and acquire second user characteristic information based on the corresponding information. In other words, first user characteristic information related to text, subject, or emotion may be acquired in response to linguistic acoustics, and second user characteristic information related to object identification information or object state information may be acquired in response to non-linguistic acoustics. That is, the user characteristic information of the present invention may include first user characteristic information and second user characteristic information acquired depending on whether the acoustic information includes linguistic acoustics or non-linguistic acoustics.
[0123] According to another embodiment of the present invention, the step of acquiring user characteristic information may include a step of acquiring user characteristic information based on audio characteristic information corresponding to the audio information. Here, the audio characteristic information may include second characteristic information related to object classification. The second characteristic information may include information regarding how many objects the audio information includes. For example, the second characteristic information corresponding to the first audio information may include information that the first audio information includes speeches of three users, and the second characteristic information corresponding to the second audio information may include information that the second audio information includes sounds related to the operation of a washing machine and sounds related to a cat's meowing.
[0124] Specifically, the server 100 may acquire user characteristic information based on the second characteristic information. For example, when the server 100 identifies that the audio information includes the speech of a large number of users through the second characteristic information, the server 100 may acquire user characteristic information that a large number of users live in a specific space. In the embodiment, the server 100 may detect a change in the number of users in a specific space corresponding to each time period through the second characteristic information of the audio information acquired for each time period, and generate user characteristic information corresponding thereto. That is, the server 100 may grasp the activity pattern or life pattern of the user in the specific space through the second characteristic information, and generate user characteristic information.
[0125] According to an embodiment of the present invention, in step S130, the server 100 may perform a step of providing matching information corresponding to the user characteristic information. For example, the matching information may be information related to advertisement. That is, providing the matching information may mean providing an efficient advertisement that enhances the user's purchasing intention, i.e., optimized advertisement information.
[0126] 6, if the user characteristic information acquired based on the sound information related to a specific space includes second user characteristic information related to the operation of a B brand washing machine 22, the server 100 can provide matching information related to a B brand dryer to the user 10 through the second user characteristic information. As another example, if the measurement space includes second user characteristic information related to the operation of a C brand air conditioner 24, the server 100 can provide matching information related to summer related products (umbrellas, travel, etc.) to the user 10 through the second user characteristic information. The above-mentioned specific descriptions of the user characteristic information are merely examples, and the present invention is not limited thereto.
[0127] According to an embodiment, the step of providing matching information corresponding to the user characteristic information may include providing the matching information based on the acquisition time and frequency of each of the plurality of user characteristic information corresponding to each of the plurality of acoustic information acquired during a predetermined time. In this case, the predetermined time may mean, for example, one day (i.e., 24 hours). In other words, the server 100 may provide the matching information based on the acquisition time and frequency of each of the plurality of user characteristic information corresponding to each of the plurality of acoustic information acquired on a 24-hour basis.
[0128] For example, when the same type of user characteristic information is continuously acquired at the same time (or the same time period) based on a 24-hour cycle, the server 100 can provide matching information corresponding to the corresponding time. For example, when user characteristic information related to the operation of a brand A washing machine is acquired at the same time (e.g., 7 p.m.) every day, the server 100 can provide matching information related to a brand A dryer at 8 p.m. when the washing is completed based on the acquired user characteristic information.
[0129] Also, for example, when a particular keyword is spoken by a user a predetermined number of times (e.g., three times) based on a 24-hour cycle, the server 100 may provide matching information corresponding to the keyword. For example, when the keyword "puppy food" is spoken by a user three or more times in a day, the server 100 may provide matching information corresponding to puppy food. The above-mentioned specific descriptions of the time and number of times of acquiring user characteristic information and the corresponding matching information are merely examples, and the present invention is not limited thereto.
[0130] That is, the server 100 records the number of times a particular keyword is spoken and the time point, and based on the record, identifies the characteristics related to the user and provides matching information through the same. In this case, since matching information suitable for the user can be provided at a more appropriate time, the advertising effect can be maximized.
[0131] According to another embodiment of the present invention, when the user characteristic information includes the first user characteristic information and the second user characteristic information, the step of providing the matching information corresponding to the user characteristic information may include the steps of acquiring related information related to the relationship between the first user characteristic information and the second user characteristic information, updating the matching information based on the related information, and providing the updated matching information. Here, the first user characteristic information may be information related to the user characteristic acquired through linguistic sound, and the second user characteristic information may be information related to the user characteristic acquired through non-linguistic sound. That is, user characteristic information corresponding to linguistic sound and non-linguistic sound corresponding to one piece of sound information may be acquired. In this case, the server 100 may update the matching information based on the related information between each acquired user characteristic information. Here, the related information may be information related to the relationship between linguistic sound and non-linguistic sound. In addition, the update of the matching information may be for further increasing the advertising effect through the matching information. For example, the update of the matching information may be related to increasing the items of the matching information to be exposed or an additional discount event applied to the matching information.
[0132] More specifically, the acoustic information may include linguistic acoustics and non-linguistic acoustics. In this case, the first user characteristic information related to the text, the subject, or the emotion may be obtained by utilizing the first acoustic model corresponding to the linguistic acoustics. In addition, the second user characteristic information related to the object identification information and the object state information may be obtained by utilizing the second acoustic model corresponding to the non-linguistic acoustics. The server 100 may obtain the association information between the first user characteristic information and the second user characteristic information. For example, the association information may be information indicating the association between each piece of user characteristic information as a numerical value. For example, if the first user characteristic information includes information that the subject is "dryer" and the second user characteristic information includes information that is related to the use of A brand dryer, the server 100 may determine that the degree of association between each piece of user characteristic information is very high and generate association information corresponding to the numerical value "98". As another example, if the first user characteristic information includes information on the subject of "dryer" and the second user characteristic information includes information related to the use of a brand B washing machine, the server 100 may determine that the degree of relevance between the user characteristic information is relatively high and generate related information corresponding to a numerical value of "85." As yet another example, if the first user characteristic information includes information on the subject of "vacuum cleaner" and the second user characteristic information includes information related to the sound of a cat meowing, the server 100 may determine that there is little relevance between the user characteristic information and generate related information corresponding to a numerical value of "7." The specific descriptions of the user characteristic information and related information described above are merely examples, and the present invention is not limited thereto.
[0133] In addition, the server 100 may update the matching information if the related information is equal to or greater than a predetermined value, and may provide the updated matching information to the user.
[0134] In a more specific example, a user utterance related to a linguistic sound (e.g., "Why is it not drying so much~", first user characteristic information) and a dryer sound, which is a non-linguistic sound (i.e., second user characteristic information), may be acquired from the sound information. In this case, the predetermined value may be 90, and the relational information between each user characteristic information may be 98. In this case, the server 100 may identify that the relational information between each user characteristic information is equal to or greater than the predetermined value, and may update the matching information. That is, when highly related linguistic sound and non-linguistic sound are simultaneously acquired through one piece of sound information, the server 100 may determine that the user has a higher interest in the corresponding object, and may update the matching information. For example, the server 100 may update the matching information to include a dryer of another company in addition to a dryer of brand A in order to provide the user with more diverse information on the corresponding object. As another example, the server 100 may update the matching information to include event information related to a discount method for purchasing a dryer of brand A. The above-mentioned specific description regarding the update of the matching information is merely exemplary, and the present invention is not limited thereto.
[0135] 5, when acquiring user characteristic information, the server 100 can identify whether the corresponding user characteristic information includes first user characteristic information and second user characteristic information. In one embodiment, when the user characteristic information acquired corresponding to the audio information includes the first user characteristic information, the server 100 can provide matching information corresponding to the first user characteristic information. As a specific example, referring to FIG. 6, when acquiring first user characteristic information that the user 10 has made an utterance related to a subject related to "puppy treats" based on the audio information (e.g., the user's utterance), the server 100 can provide matching information related to "puppy treats". The above-mentioned specific description of the first user characteristic information and the matching information is merely an example, and the present invention is not limited thereto.
[0136] In addition, when the user characteristic information acquired corresponding to the sound information includes second user characteristic information, the server 100 may provide matching information corresponding to the second user characteristic information. For a specific example, referring to FIG. 6, when second user characteristic information is acquired that a puppy 23 is located in a specific space as a companion animal based on sound information (e.g., a puppy's barking), the server 100 may provide various matching information related to the puppy, such as puppy treats, puppy feed, puppy toys, or puppy clothes. For another example, when second user characteristic information is acquired that a user is not in good health based on sound information (e.g., a user's coughing), the server 100 may provide various matching information related to the user's health promotion, such as cold medicine, porridge, tea, or health supplements. The above-mentioned specific description of the second user characteristic information and matching information is merely illustrative, and the present invention is not limited thereto.
[0137] In addition, when the user characteristic information acquired corresponding to the audio information includes the first user characteristic information and the second user characteristic information, i.e., all of them, the server 100 can acquire association information between each piece of user characteristic information and update the matching information based on the association information. Also, the server 100 can provide the updated matching information.
[0138] In other words, when relevant verbal and non-verbal voices are simultaneously acquired, the user's interest is higher, so the server 100 can provide matching information reflecting a large amount of information for decision-making or event information related to additional discounts related to the purchase of an object in order to increase the user's purchase possibility. That is, it is possible to provide optimized matching information by predicting the user's interest based on the related information of verbal and non-verbal voices in the acquired audio information and providing the corresponding matching information. This provides differentiated matching information according to the user's interest, and has the effect of maximizing the possibility of purchase conversion.
[0139] According to an embodiment of the present invention, the step of providing matching information based on the user characteristic information may include the steps of generating an environmental characteristic table based on one or more user characteristic information corresponding to each of one or more pieces of acoustic information acquired at a predetermined time period, and providing the matching information based on the environmental characteristic table. The environmental characteristic table may be information on statistics of each piece of user characteristic information acquired at a predetermined time period. In one embodiment, the predetermined time period may mean 24 hours. In other words, the server 100 can generate an environmental characteristic table related to the statistical values of user characteristic information acquired for each time period based on a 24-hour time period, and provide matching information based on the corresponding environmental characteristic table.
[0140] As a specific example, by observing the statistics of user characteristic information acquired by time period through the environment characteristic table, the server 100 can provide matching information related to food or interior props to a user who is very active at home. Also, for example, the server 100 can provide matching information related to new viewing content to a user who spends a lot of time watching TV 21 at home. As another example, the server 100 can provide matching information related to health supplements or unmanned laundry services to a user who is relatively inactive at home.
[0141] In addition, in one embodiment, the step of providing matching information corresponding to the user characteristic information may include identifying a first time point for providing matching information based on an environmental characteristic table, and providing matching information corresponding to the first time point. The first time point may mean an optimal time point for providing matching information to the corresponding user. That is, the user's activity routine in a specific space may be grasped through the environmental characteristic table, and optimal matching information for each time point may be provided. For example, if the sound of a washing machine is periodically identified during a first time period (8 p.m.) through the environmental characteristics table, matching information related to a fabric softener or a dryer, etc., can be provided during the first time period.
[0142] As another example, if the sound of a vacuum cleaner is periodically identified during a second time period (2 p.m.) through the environmental characteristics table, matching information related to a wireless vacuum cleaner or a wet rag vacuum cleaner during the second time period can be provided. That is, the server 100 can maximize the effectiveness of advertisement by providing appropriate matching information in relation to a specific activity point in time. The server 100 can maximize the effectiveness of advertisement by providing targeted matching information based on the acoustic information acquired in relation to the user's living environment according to various embodiments as described above.
[0143] The steps of a method or algorithm described in connection with the embodiments of the present invention may be embodied directly in hardware, in a software module executed by hardware, or in a combination thereof. The software module may reside in a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable ROM (EPROM), an Electrically Erasable Programmable ROM (EEPROM), a Flash Memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer readable recording medium commonly known in the art to which the present invention pertains.
[0144] The components of the present invention may be embodied as a program (or application) and stored on a medium for execution in conjunction with a computer, which is hardware. The components of the present invention may be implemented as software programming or software elements, and similarly, embodiments include various algorithms embodied as a combination of data structures, processes, routines, or other programming constructs, and may be embodied in programming or scripting languages such as C, C++, Java, assembler, etc. Functional aspects may be embodied as algorithms executed by one or more processors.
[0145] Although the embodiments of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention can be embodied in other specific forms without changing the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.
[0146] The above description is of the best mode for carrying out the invention. [Industrial Applicability]
[0147] The present invention can be utilized in the field of services that provide matching information through audio information analysis.
Claims
1. A method performed on a computing device, comprising: acquiring acoustic information; acquiring user characteristic information based on the acoustic information; and providing matching information corresponding to the user characteristic information.
2. The step of acquiring user characteristic information based on the acoustic information includes: identifying a user or an object through analysis of the acoustic information; and The method of claim 1 , further comprising: generating the user characteristic information based on the identified user or object.
3. The step of acquiring user characteristic information based on the acoustic information includes: generating activity time information related to a time during which a user is active in a specific space through analysis of the acoustic information; and The method of claim 1 , further comprising: generating the user characteristic information based on the activity time information.
4. The step of providing matching information corresponding to the user characteristic information includes:
2. The method of claim 1, further comprising: providing the matching information based on the acquisition time and frequency of each of a plurality of pieces of user characteristic information corresponding to each of a plurality of pieces of acoustic information acquired during a predetermined time period.
5. The step of acquiring acoustic information includes: performing pre-processing on the acquired acoustic information; and identifying acoustic feature information corresponding to the pre-processed acoustic information; The acoustic characteristic information is The method of claim 1, further comprising: first characteristic information regarding whether the acoustic information is related to at least one of linguistic acoustics and non-linguistic acoustics; and second characteristic information regarding an object classification.
6. The step of acquiring user characteristic information based on the acoustic information includes: acquiring the user characteristic information based on the acoustic characteristic information corresponding to the acoustic information; The step of acquiring user characteristic information includes: If the acoustic feature information includes first feature information related to linguistic sounds, processing the acoustic information as an input to a first acoustic model to obtain user feature information corresponding to the acoustic information; or and if the acoustic characteristic information includes first characteristic information corresponding to non-linguistic acoustics, processing the acoustic information as an input to a second acoustic model to obtain user characteristic information corresponding to the acoustic information.
7. The first acoustic model includes a neural network model trained to perform an analysis on acoustic information related to linguistic acoustics and to identify at least one of a text, a theme, or an emotion associated with the acoustic information, The second acoustic model includes a neural network model trained to perform an analysis on acoustic information related to a non-linguistic sound and obtain object identification information or object state information related to the acoustic information, The user characteristic information is The method of claim 6, further comprising: first user characteristic information related to at least one of text, subject or emotion related to the acoustic information; and second user characteristic information related to the object identification information or the object state information related to the acoustic information.
8. The step of providing matching information corresponding to the user characteristic information includes: acquiring association information relating to an association between the first user characteristic information and the second user characteristic information, when the acquired user characteristic information includes the first user characteristic information and the second user characteristic information; updating matching information based on the related information; and The method of claim 7 , further comprising: providing the updated matching information.
9. The step of providing matching information based on the user characteristic information includes: generating an environmental characteristic table based on one or more pieces of user characteristic information corresponding to one or more pieces of acoustic information acquired at a predetermined time interval; and providing the matching information based on the environmental characteristic table; The environmental characteristics table includes: The method of claim 1 , wherein the information is information on statistics of each of the user characteristic information acquired at predetermined time intervals.
10. The step of providing matching information corresponding to the user characteristic information includes: identifying a first time point for providing the matching information based on the environmental characteristic table; and 10. The method of claim 9, further comprising: providing the matching information corresponding to the first time point.
11. memory storing one or more instructions; and An apparatus comprising a processor for executing the one or more instructions stored in the memory, the processor performing the method of claim 1 by executing the one or more instructions.
12. A computer program stored on a computer-readable recording medium so as to be combined with a computer that is hardware and capable of carrying out the method according to claim 1.