Methods, apparatus, and procedures for providing matching information by analyzing sound information.
Patent Information
- Application Number
- CN202280035411.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-23
- Filing Date
- 2022-04-07
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2042-04-07
AI Technical Summary
并且,随着基于预设用户兴趣领域的用户信息提供在线广告,当用户的兴趣领域产生变化时,除非用户直接改变所设定的兴趣领域,否则仅显示与先前设定的兴趣领域相关的广告,因此,将无法提供用户的新兴趣领域信息
[0028]根据本发明多个实施例,本发明可基于所获取的与用户生活环境相关的声音信息提供目标匹配信息来最大限度地提高广告效果。
Smart Images

Figure CN117678016B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for providing appropriate matching information to a user, and more specifically, to a technique for providing optimal matching information to a user by analyzing sound information. Background Technology
[0002] With the increasing use of various electronic devices such as smart TVs, smartphones, and tablets (PCs) and the widespread availability of internet services, advertising delivered through electronic devices or online is gradually increasing.
[0003] For example, methods of delivering advertising through electronic devices or online include displaying advertisements set by advertisers on various websites as banners to all users visiting those websites. As a specific example, with the recent trend of setting target audiences for online advertising, customized advertisements are being delivered to these target audiences.
[0004] Methods for targeting ad viewers to deliver customized advertising include obtaining visitor information online and identifying visitor areas of interest through analysis.
[0005] This type of online advertising, as a limited form of advertising, collects information from visitors when they access a specific website and use its services. Furthermore, since online advertising is provided based on user information within preset user interest areas, when a user's interest areas change, unless the user directly changes their set interest areas, only ads related to the previously set interest areas will be displayed; therefore, information about the user's new interest areas will not be provided.
[0006] Therefore, the existing advertising methods described above are insufficient to improve advertising efficiency in terms of delivering ads based on visitors' areas of interest. Furthermore, these methods cannot proactively respond to users' changing interests over time.
[0007] Existing technical documents
[0008] Patent documents
[0009] Korean Patent 10-2044555 Summary of the Invention
[0010] Technical issues
[0011] To address the aforementioned problems, the purpose of this invention is to provide users with more appropriate matching information by analyzing sound information.
[0012] The objectives of this invention are not limited to those mentioned above. Those skilled in the art to which this invention pertains can clearly understand other objectives not mentioned through the following description.
[0013] Technical solution
[0014] To achieve the above objectives, several embodiments of the present invention disclose a method for providing matching information by analyzing sound information. The method may include the following steps: acquiring sound information; acquiring user characteristic information based on the sound information; and providing matching information corresponding to the user characteristic information.
[0015] In an alternative embodiment, the step of obtaining user characteristic information based on the above-mentioned sound information may include the following steps: identifying a user or object by analyzing the above-mentioned sound information; and generating the above-mentioned user characteristic information based on the identified user or object.
[0016] In an alternative embodiment, the step of obtaining user characteristic information based on the above-mentioned sound information may include the following steps: generating activity time information related to the time of the user's activity in a specific space by analyzing the above-mentioned sound information; and generating the above-mentioned user characteristic information based on the above-mentioned activity time information.
[0017] In an alternative embodiment, the step of providing matching information corresponding to the aforementioned user characteristic information may include the following steps: providing the matching information based on the acquisition time points and frequencies of multiple user characteristic information corresponding to multiple voice information acquired within a predetermined time period.
[0018] In an alternative embodiment, the above-described steps of acquiring sound information include the following steps: performing preprocessing on the acquired sound information; and identifying sound characteristic information corresponding to the preprocessed sound information, wherein the sound characteristic information may include first characteristic information and second characteristic information, wherein the first characteristic information is used to determine whether the sound information is related to at least one of speech sounds and non-speech sounds, and the second characteristic information is used to distinguish objects.
[0019] In an alternative embodiment, the step of obtaining user characteristic information based on the above-mentioned sound information includes the following steps: obtaining the user characteristic information based on the sound characteristic information corresponding to the above-mentioned sound information, and the step of obtaining the user characteristic information may include at least one of the following steps: when the above-mentioned sound characteristic information contains first characteristic information related to speech sounds, inputting the above-mentioned sound information into a first sound model to obtain user characteristic information corresponding to the above-mentioned sound information; or when the above-mentioned sound characteristic information contains first characteristic information related to non-speech sounds, inputting the above-mentioned sound information into a second sound model to obtain user characteristic information corresponding to the above-mentioned sound information.
[0020] In an alternative embodiment, the first sound model includes a neural network model that learns to identify at least one of text, topic, or emotion related to the sound information by performing analysis on sound information related to speech sounds, and the second sound model includes a neural network model that learns to obtain object recognition information or object state information related to the sound information by performing analysis on sound information related to non-speech sounds, and the user characteristic information may include at least one of first user characteristic information and second user characteristic information, wherein the first user characteristic information is at least one of text, topic, or emotion related to the sound information, and the second user characteristic information is the object recognition information or object state information related to the sound information.
[0021] In an alternative embodiment, the step of providing matching information corresponding to the aforementioned user characteristic information may include the following steps: when the acquired user characteristic information includes the aforementioned first user characteristic information and the aforementioned second user characteristic information, obtaining association information regarding the correlation between the aforementioned first user characteristic information and the aforementioned second user characteristic information; updating the matching information based on the aforementioned association information; and providing the updated matching information.
[0022] In an alternative embodiment, the step of providing matching information based on the aforementioned user characteristic information includes the following steps: generating an environmental characteristic list based on one or more user characteristic information corresponding to one or more voice information acquired according to a predetermined time period; and providing the matching information based on the aforementioned environmental characteristic list, wherein the environmental characteristic list may be information that statistically analyzes multiple user characteristic information acquired according to the aforementioned predetermined time period.
[0023] In an alternative embodiment, the step of providing matching information corresponding to the aforementioned user characteristic information may include the following steps: identifying a first time point for providing the aforementioned matching information based on the aforementioned environmental characteristic list; and providing the aforementioned matching information corresponding to the aforementioned first time point.
[0024] A further embodiment of the present invention discloses an apparatus for performing a method of providing matching information by analyzing sound information. The apparatus may include: a memory for storing one or more instructions; and a processor for executing the one or more instructions stored in the memory, wherein the processor can execute the method of providing matching information by analyzing sound information by executing the one or more instructions.
[0025] Another embodiment of the present invention discloses a computer program stored in a computer-readable recording medium. This computer program can be combined with a computer as hardware to execute the method described above for providing matching information by analyzing sound information.
[0026] Other specific aspects of the present invention can be found in the following detailed description and accompanying drawings.
[0027] The effects of the invention
[0028] According to various embodiments of the present invention, the present invention can provide target matching information based on the acquired sound information related to the user's living environment to maximize the advertising effect.
[0029] The effects of this invention are not limited to those mentioned above. Those skilled in the art to which this invention pertains can clearly understand other effects not mentioned through the following description. Attached Figure Description
[0030] Figure 1 A system diagram is provided for briefly illustrating a method for providing matching information by analyzing sound information to perform an embodiment of the present invention.
[0031] Figure 2 This is a hardware structure diagram of a server used to provide a method for providing matching information by analyzing sound information according to an embodiment of the present invention.
[0032] Figure 3 The flowchart illustrates a method for providing matching information by analyzing sound information according to an embodiment of the present invention.
[0033] Figure 4 The following is a flowchart illustrating an embodiment of the present invention for obtaining user characteristic information based on sound information.
[0034] Figure 5 The following is a flowchart illustrating an embodiment of the present invention for providing matching information based on user characteristic information.
[0035] Figure 6 This is an illustrative diagram used to illustrate an embodiment of the present invention of a process for acquiring various types of sound information within a user positioning space and a method for providing matching information corresponding to the sound information.
[0036] Best practice
[0037] Hereinafter, several embodiments are described with reference to the accompanying drawings. In this specification, the various embodiments are only used to understand the present invention. However, it should be understood that such embodiments may also be practiced without specific description.
[0038] In this specification, the terms "component," "module," "system," etc., as used, refer to computer-related entities, hardware, firmware, software, combinations of software and hardware, or the execution of software. For example, a component can be a process, process, object, thread of execution, program, and / or computer running on a processor, but is not limited thereto. For example, an application program executed by a computing device and the computing device itself can both be components. More than one component may exist within a processor and / or thread of execution. A component may reside within a single computer, or a component may be distributed among two or more computers. Furthermore, such a component may be executed by multiple computer-readable media having various data structures stored within it. For example, multiple components may communicate through local and / or remote processing based on signals having more than one data packet (e.g., receiving data and / or signals from a component interacting with other components in a local system or distributed system and transmitting data to other systems via a network such as the Internet).
[0039] Furthermore, the term "or" refers to an inclusive "or" rather than an exclusive "or." That is, unless otherwise defined or explicitly stated in the context, "X uses A or B" refers to one meaning within a natural inclusive substitution. Specifically, it can represent all cases such as "X uses A or X uses B," "X uses both X and B," and "X uses A or B." Moreover, in this specification, the term "and / or" refers to all combinations including more than one of the related items.
[0040] Furthermore, terms such as "comprising" and / or "including" refer to the presence of the corresponding feature and / or structural element. However, the terms "comprising" and / or "including" do not exclude the possibility of the presence or addition of more than one other feature, structural element, and / or combination thereof. And, unless otherwise defined in this specification and claims or expressly indicated otherwise in the context, the singular expression means "one or more".
[0041] Those skilled in the art to which this invention pertains will understand that, in relation to the embodiments disclosed herein, the various exemplary logic blocks, structures, modules, circuits, devices, logic, and algorithm steps described can be implemented by electronic hardware, computer software, or a combination thereof. Hereinafter, various exemplary components, blocks, structures, devices, logic, modules, circuits, and steps will be described according to their functions to clearly illustrate the interchangeability of hardware and software. However, whether a function is implemented by hardware or software depends on the specific application and design constraints imposed on the overall system. Those skilled in the art to which this invention pertains can implement the functions described by various methods for each specific application. However, such a decision on implementation should not be construed as departing from the scope of this invention.
[0042] The disclosed embodiments are for illustrative purposes only, enabling those skilled in the art to readily utilize or implement the invention. Various modifications can be made to the embodiments by those skilled in the art. The general principles defined herein can be applied to other embodiments without departing from the scope of the invention. Therefore, the invention is not limited to the embodiments disclosed herein. The invention should be interpreted based on the principles disclosed herein and the widest scope consistent with novel features.
[0043] In this specification, "computer" refers to all types of hardware devices including at least one processor. According to embodiments of the invention, it should be understood to include the meaning of the corresponding hardware device running software structures. For example, a computer may include smartphones, tablets, desktop computers, laptops, and user clients and applications driven by each device, but is not limited thereto.
[0044] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0045] The steps described in this specification can be performed by a computer; however, the subject matter of each step is not limited thereto. According to embodiments of the invention, at least a portion of each step can also be performed by different devices.
[0046] According to various embodiments of the present invention, the method for providing matching information by analyzing sound information can provide the most appropriate matching information to each user based on acquiring various sound information from multiple users in their real lives. For example, the matching information can be advertising-related information. That is, providing the most appropriate matching information to a user means providing the user with advertisements that effectively increase their purchase desire, i.e., providing the most appropriate advertising information. In other words, the method of providing matching information by analyzing sound information of the present invention can analyze various sound information acquired from a user's living space to provide customized advertising information to the corresponding user. From the advertiser's perspective, this allows for selective display of advertisements to potential or target customers interested in the advertisement, thus significantly reducing advertising costs and maximizing advertising effectiveness. Furthermore, from the consumer's perspective, since they only receive advertisements that interest them or satisfy their needs, it increases the convenience of information retrieval.
[0047] Figure 1 A system diagram is provided to briefly illustrate a method for providing matching information by analyzing sound information in carrying out an embodiment of the present invention. (See diagram below.) Figure 1 As shown, in order to perform the method of providing matching information by analyzing sound information, a system according to an embodiment of the present invention may include a server 100, a user terminal 200, and an external server 300.
[0048] in, Figure 1The system shown that provides matching information by analyzing sound information is only one embodiment of the present invention, and its structural elements are not limited thereto. Its structural elements may be added, changed or omitted as needed.
[0049] In one embodiment of the present invention, a server 100 that provides matching information by analyzing sound information can acquire sound information and provide the most appropriate matching information by analyzing the acquired sound information. That is, the server 100 that provides matching information by analyzing sound information acquires various sound information of users in real life, and by analyzing the acquired sound information through a sound model to identify information related to the user's interests, the most appropriate matching information can be provided to the corresponding user.
[0050] According to embodiments of the present invention, the server 100, which provides matching information by analyzing sound information, may include any server implemented based on an application programming interface (API). For example, a user terminal 200 acquires sound information and performs analysis on it, and can provide sound information recognition results to the server 100 through the API. The sound information recognition results refer to features related to the sound information analysis. As a specific example, the sound information recognition results can be a spectrogram obtained by processing the sound information using a Short-Time Fourier Transform (STFT). The spectrogram is used to visualize and identify sounds or vibrations and can be composed of a combination of waveform and spectrum features. The spectrogram can represent amplitude differences as differences in printing density or display color based on changes in the time and frequency axes.
[0051] As another example, the sound information recognition result may include a Mel-spectrum obtained by processing the spectrogram using a Mel-filter bank. Typically, the vibrational components of the human cochlea vary with the frequency of the speech data. The human cochlea is adept at detecting frequency changes in the low-frequency band but not so good at detecting frequency changes in the high-frequency band. Therefore, a Mel-filter bank can be applied to obtain a Mel-spectrum from the spectrogram to achieve a recognition capability similar to the characteristics of the human cochlea in detecting speech data. That is, a small number of filters are applied in the low-frequency band, and a wider range of filters can be gradually applied in the high-frequency band. In other words, in order to recognize sound information with similar characteristics to the human cochlea, the user terminal 200 may apply a Mel-filter bank to the spectrogram to obtain a Mel-spectrum. That is, the Mel-spectrum may include frequency components reflecting human auditory characteristics.
[0052] Server 100 can provide the most appropriate matching information to the corresponding user based on the voice information recognition results obtained from user terminal 200. In this case, server 100 receives voice information recognition results (e.g., spectrograms or Mel spectrograms based on user interest prediction) from user terminal 200, thus resolving user privacy issues arising from the collection of voice information.
[0053] In one embodiment of the invention, the sound model (e.g., an artificial intelligence model) consists of one or more network functions, typically composed of a set of interconnected computational units that can be called "nodes." These "nodes" can also be called "neurons." One or more network functions include at least one node. The nodes (or neurons) constituting one or more network functions can be connected through one or more "links."
[0054] Within an artificial intelligence model, one or more nodes connected by links can form relative relationships between input and output nodes. Input and output nodes are relative concepts; any node that is an output node relative to one node can also be an input node relative to another node, and vice versa. As mentioned above, relationships between input and output nodes can be generated around links. One or more output nodes can be connected to an input node through a link, and vice versa.
[0055] In a relationship between input and output nodes connected by a link, the output node can determine its value based on the data input to the input nodes. The nodes connecting the input and output nodes can have weighted values. These weighted values are variable and can be changed by the user or algorithm to enable the AI model to perform the desired function. For example, when more than one input node is connected to an output node via a link, the output node can determine its value based on multiple input values from the input nodes connected to the output node and the weighted values set by the links between the corresponding input nodes.
[0056] As mentioned above, in an artificial intelligence model, one or more nodes are connected by one or more links, forming input and output node relationships within the model. The characteristics of an artificial intelligence model can be determined based on the number of nodes and links, the connection relationships between nodes and links, and the weighted values assigned to the links. For example, when two artificial intelligence models exist with the same number of nodes and links but different weighted values among multiple links, the two models can be identified as distinct.
[0057] In an artificial intelligence model, a subset of nodes can form a layer based on their distance from the initial input node. For example, a set of nodes at a distance of n from the initial input node can form n layers. The distance from the initial input node can be defined as the minimum number of links required to reach the corresponding node from the initial input node. However, such layers are randomly defined for illustrative purposes, and the order of layers within the artificial intelligence model can be defined using a different method. For example, a node's layer can also be defined as its distance from the final output node.
[0058] In the relationship between multiple nodes within an artificial intelligence model and other nodes, the initial input node refers to one or more nodes into which data is directly input without being linked. Alternatively, in the relationship between nodes within the artificial intelligence model network based on links, it refers to nodes that are not present in other input nodes connected by links. Similarly, in the relationship between multiple nodes within an artificial intelligence model and other nodes, the final output node refers to one or more nodes that do not have an output node. Furthermore, hidden nodes refer to the nodes that constitute the artificial intelligence model, not the initial input node or the final output node. In an artificial intelligence model according to an embodiment of the present invention, the number of nodes in the input layer may be greater than the number of nodes in the hidden layer near the output layer, and it can be an artificial intelligence model in which the number of nodes decreases as the model progresses from the input layer to the hidden layer.
[0059] Artificial intelligence models can include more than one hidden layer. Hidden nodes in a hidden layer can take the outputs of previous layers and surrounding hidden nodes as input. The number of hidden nodes in each hidden layer can be the same or different. The number of nodes in the input layer depends on the number of data fields in the input data and can be the same as or different from the number of hidden nodes. Input data input to the input layer can be processed by the hidden nodes of the hidden layer and output through a fully connected layer (FCL) as the output layer.
[0060] In several embodiments of the present invention, the artificial intelligence model can achieve supervised learning by treating multiple sound information and specific information corresponding to each sound information as learning data. However, it is not limited to this, and various learning methods can also be applied.
[0061] Supervised learning, which typically generates learning data by labeling specific data and related information, refers to a method of learning through the generated learning data.
[0062] In one embodiment of the present invention, when the learning of one or more network functions has been performed in a predetermined batch or more, the server 100, which provides matching information by analyzing sound information, can use verification data to determine whether to terminate the learning. The predetermined batch can be part of the overall learning target batch.
[0063] The verification data may consist of at least a labeled portion of the learning data. That is, the server 100, which provides matching information by analyzing sound information, executes the learning of an artificial intelligence model using the learning data. After the learning of the artificial intelligence model is repeatedly executed in predetermined batches, the verification data can be used to determine whether the learning effect of the artificial intelligence model has reached or exceeded a predetermined level. For example, when using 100 pieces of learning data to perform 10 iterations of learning, after performing 10 iterations of learning in predetermined batches, the server 100, which provides matching information by analyzing sound information, can use 10 pieces of verification data to perform 3 iterations of learning. During these 3 iterations, if the output of the artificial intelligence model falls below a predetermined level, it can be determined that further learning is meaningless and the learning process ends.
[0064] That is, the validation data can be used to determine whether the learning of each batch has reached or fallen below a specified effect during the iterative learning of the artificial intelligence model, thus determining whether the learning is complete. The quantity of learning data and validation data and the number of iterations mentioned above are merely examples, and the present invention is not limited thereto.
[0065] A server 100, which provides matching information by analyzing sound information, determines whether to activate one or more network functions based on test data, thereby generating an artificial intelligence model. The test data, which can be used to verify the performance of the artificial intelligence model, may consist of at least a portion of the learning data. For example, 70% of the learning data can be used for learning the artificial intelligence model (i.e., for adjusting weights to output results similar to the labels), and 30% of the learning data can be used as test data to verify the performance of the artificial intelligence model. The server 100, which provides matching information by analyzing sound information, inputs the test data into the learned artificial intelligence model and measures the error. Whether to activate the artificial intelligence model is determined based on whether the measured error reaches a predetermined performance level.
[0066] For an AI model that has completed learning, a server 100 that provides matching information by analyzing sound information can use test data to verify the performance of the AI model that has completed learning. If the AI model that has completed learning reaches a performance level above a predetermined level, the corresponding AI model can be activated for use in other applications.
[0067] Furthermore, if the learned artificial intelligence model achieves performance below a predetermined level, the server 100, which provides matching information by analyzing sound information, can deregister and delete the corresponding artificial intelligence model. For example, the optimal stimulus location calculation server 100 can judge the performance of the generated artificial intelligence model based on factors such as accuracy, precision, and recall. The above performance evaluation benchmarks are merely examples, and the present invention is not limited thereto. According to one embodiment of the present invention, the optimal stimulus location calculation server 100 can independently learn various artificial intelligence models and generate multiple artificial intelligence models, and can use only artificial intelligence models with performance exceeding a specified level by evaluating their performance. However, it is not limited thereto.
[0068] In this specification, the terms "operational model," "neural network," "network function," and "neural network" can be used interchangeably (hereinafter collectively referred to as "neural network"). A data structure may include a neural network. Furthermore, a data structure including a neural network may be stored on a computer-readable medium. A data structure including a neural network may include data input to the neural network, weights of the neural network, hyperparameters of the neural network, data obtained from the neural network, activation functions associated with nodes or layers of the neural network, and a loss function used to train the neural network. A data structure including a neural network may include any structural element of the above-disclosed structures. That is, a data structure including a neural network may include data input to the neural network, weights of the neural network, hyperparameters of the neural network, data obtained from the neural network, activation functions associated with nodes or layers of the neural network, a loss function used to train the neural network, etc., or any combination thereof. In addition to the structures described above, a data structure including a neural network may include any other information used to determine the characteristics of the neural network. Furthermore, the data structure may include all types of data generated from or used in the computation of the neural network, and is not limited to the above. A computer-readable medium may include a computer-readable recording medium and / or a computer-readable transmission medium. A neural network may consist of a collection of interconnected computational units commonly referred to as nodes. Such nodes may also be called neurons. A neural network includes at least one node.
[0069] According to one embodiment of the present invention, the server 100 that provides matching information by analyzing voice information can be a server providing cloud computing services. More specifically, the server 100 that provides matching information by analyzing voice information can be a server providing cloud computing services, which is a computer that processes data through another computer connected to the Internet, rather than the user's computer, based on Internet computing. The aforementioned cloud computing service can be a service that stores data on the Internet, and can be used anywhere by accessing the Internet even without setting up the user's required data or programs on its own computer, and can easily share and transmit data stored on the Internet through simple operations and clicks. Furthermore, the cloud computing service not only stores data on a server on the Internet, but also allows the desired work to be performed through the application functions provided by the webpage without setting up separate programs, meaning that multiple people can share files and perform work simultaneously. Furthermore, the cloud computing service can be implemented by at least one of Infrastructure as a Service (IaaS), Platform as a Service (PaaS), Software as a Service (SaaS), virtual machine cloud servers, and container cloud servers. That is, the service 100 of the present invention that provides matching information by analyzing voice information can be implemented by at least one of the above-mentioned cloud computing services. The cloud computing services described above are merely examples and may also include any platform used to construct the cloud computing environment of this invention.
[0070] In several embodiments of the present invention, a server 100 that provides matching information by analyzing sound information can be connected to a user terminal 200 via a network 400. It can generate a sound model for analyzing sound information. In addition, it can provide optimal matching information for each user based on information (e.g., user characteristic information) obtained by analyzing sound information through the sound model.
[0071] In this context, Network 400 refers to a connection structure capable of exchanging information with multiple terminals and servers. For example, Network 400 may include Local Area Network (LAN), Wide Area Network (WAN), World Wide Web (WWW), wired and wireless data communication networks, telephone networks, and wired and wireless television communication networks.
[0072] Furthermore, wireless data communication networks include 3G, 4G, 5G, the 3rd Generation Partnership Project (3GPP), the 5th Generation Partnership Project (5GPP), Long Term Evolution (LTE), World Interoperability for Microwave Access (WIMAX), Wi-Fi, the Internet, Local Area Networks (LANs), Wireless Local Area Networks (WLANs), Wide Area Networks (WANs), Personal Area Networks (PANs), Radio Frequency (RF), Bluetooth networks, Near-Field Communication (NFC) networks, satellite broadcasting networks, analog broadcasting networks, and Digital Multimedia Broadcasting (DMB) networks, but are not limited to these.
[0073] In one embodiment of the present invention, a user terminal 200 can be connected to a server 100 that provides matching information by analyzing sound information via a network 400. The server 100 can provide multiple sound information (e.g., speech sound information or non-speech sound information) and can receive various information corresponding to the provided sound information (e.g., user characteristic information corresponding to the sound information and matching information corresponding to the user characteristic information).
[0074] The user terminal 200, as a wireless communication device ensuring portability and mobility, may include, but is not limited to, all types of handheld wireless communication devices such as navigators, Personal Communication Systems (PCS), Global System for Mobile Communications (GSM), Personal Digital Cellular Systems (PDC), Personal Handyphone Systems (PHS), Personal Digital Assistants (PDAs), International Mobile Telecommunications System 2000 (IMT), Code Division Multiple Access 2000 (CDMA), Wideband Code Division Multiple Access (W-CDMA), Wireless Broadband Internet (Wibro) terminals, smartphones, smartpads, and tablet PCs. For example, the user terminal 200 may also include artificial intelligence (AI) speakers and AI televisions that interact with the user based on hot words to provide various functions such as music enjoyment and information retrieval.
[0075] In one embodiment of the present invention, user terminal 200 may include a first user terminal 210 and a second user terminal 220. The user terminals (first user terminal 210 and second user terminal 220) have mechanisms for communicating with each other or with other entities via network 400. In a system that provides matching information by analyzing voice information, this refers to entities of any form. For example, the first user terminal 210 may include any terminal associated with a user receiving matching information. Furthermore, the second user terminal 220 may include any terminal associated with an advertiser registering matching information. This user terminal 200 includes a display, thus, upon receiving user input, it can provide output of any form to the user.
[0076] In one embodiment of the present invention, an external server 300 can be connected to a server 100 that provides matching information by analyzing sound information via a network 400. The server 100, which provides matching information by analyzing sound information, applies an artificial intelligence model to provide various information / data required for analyzing the sound information. Alternatively, as the artificial intelligence model performs sound information analysis, it can receive, store, and manage the exported result data. For example, the external server 300 can be a storage server, separately located outside the server 100 that provides matching information by analyzing sound information, but it is not limited to this. Referring hereafter... Figure 2 The hardware structure of server 100, which provides matching information by analyzing sound information, is described.
[0077] Figure 2 This is a hardware structure diagram of a server used to provide a method for providing matching information by analyzing sound information according to an embodiment of the present invention.
[0078] Reference Figure 2 The optimal stimulation position calculation server 100 (hereinafter referred to as "server 100") according to an embodiment of the present invention may include: one or more processors 110; a memory 120 for loading a computer program 151 executed by the processor 110; a bus 130; a communication interface 140; and an auxiliary memory 150 for storing the computer program. Figure 2 Only structural elements relevant to embodiments of the present invention are shown. Therefore, those skilled in the art to which this invention pertains should understand that, in addition to... Figure 2 In addition to the structural elements shown, other general structural elements may also be included.
[0079] The processor 110 is used to control the overall operation of the various structures of the server 100. The processor 110 may include a central processing unit (CPU), a microprocessor (MPU), a micro controller (MCU), a graphics processing unit (GPU), or any type of processor known in the art.
[0080] The processor 110 can read a computer program stored in the memory 120 and execute data processing for an artificial intelligence model according to an embodiment of the present invention. According to an embodiment of the present invention, the processor 110 can perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), extracting features from the input data, calculating errors, and updating the weights of the neural network using backpropagation.
[0081] Furthermore, processor 110 enables at least one of the central processing unit (CPU), general-purpose graphics processing unit (GPGPU), and tensor processor (TPU) to process the learning of network functions. For example, the CPU and GPGPU can jointly process the learning of network functions and the data classification using the network functions. In one embodiment of the invention, processors of multiple computing devices can jointly process the learning of network functions and the data classification using the network functions. In another embodiment of the invention, the computer program executed by the computing device can be a program executable by the CPU, GPGPU, and tensor processor.
[0082] In this specification, network functions can be used for the exchange between artificial neural networks and neural networks. In this invention, a network function may include more than one neural network, and in this case, the output of the network function can be an ensemble of the outputs of more than one neural network.
[0083] Processor 110 can provide the sound model of the present invention by reading a computer program stored in memory 120. According to one embodiment of the present invention, processor 110 can acquire user characteristic information corresponding to sound information. According to one embodiment of the present invention, processor 110 can perform calculations for learning the sound model.
[0084] According to one embodiment of the present invention, processor 110 can handle the overall operation of server 100. Processor 110 can provide or process appropriate information or functions to users or user terminals by processing signals, data, information, etc. input or output by the above-mentioned structural elements or driving applications stored in memory 120.
[0085] Furthermore, the processor 110 may perform operations of at least one application or program in order to execute the method of the embodiments of the present invention, and the server 100 may include more than one processor.
[0086] In various embodiments of the present invention, the processor 110 may further include random access memory (RAM, not shown) and read-only memory (ROM, not shown) for temporarily storing and / or permanently storing signals (or data) processed within the processor 110. Furthermore, the processor 110 may be a system-on-chip (SoC), including at least one of a graphics processor, random access memory, and read-only memory.
[0087] Memory 120 is used to store various data, instructions, and / or information. Memory 120 can load a computer program 151 from auxiliary memory 150 to perform the methods / operations of various embodiments of the present invention. If computer program 151 is loaded into memory 120, processor 110 can perform the above-described methods / operations by executing one or more instructions constituting computer program 151. Memory 120 can be a volatile memory such as random access memory; however, the scope of the present invention is not limited thereto.
[0088] Bus 130 is used to provide communication functions between the structural elements of server 100. Bus 130 can be various types of buses such as address bus, data bus, and control bus.
[0089] The communication interface 140 is used to support wireless and wired network communication of the server 100. Furthermore, the communication interface 140 can also support various communication methods besides networks. Therefore, the communication interface 140 may include communication modules known in the art. In several embodiments of the present invention, the communication interface 140 may also be omitted.
[0090] The auxiliary storage 150 can permanently store the computer program 151. When the server 100 executes a process of providing matching information by analyzing sound information, the auxiliary storage 150 can store various necessary information for providing matching information by analyzing sound information.
[0091] The auxiliary storage 150 may include a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory or other non-volatile memory, a hard disk, a portable hard disk, or any type of computer-readable recording medium known in the art to which this invention pertains.
[0092] Computer program 151 may include one or more instructions that, when loaded into memory 120, cause processor 110 to execute the methods / operations of various embodiments of the present invention. That is, processor 110 can execute the methods / operations of various embodiments of the present invention by executing the one or more instructions described above.
[0093] In one embodiment of the present invention, computer program 151 may include one or more instructions to execute a method for providing matching information by analyzing sound information, the method including the following steps: acquiring sound information; acquiring user characteristic information based on the sound information; and providing matching information corresponding to the user characteristic information.
[0094] The methods or algorithm steps of the embodiments of the present invention can be implemented directly by hardware, or by software modules executed by hardware, or by a combination thereof. The software modules can reside in random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, hard disk, portable disk, read-only optical disc (CD-ROM), or any type of computer-readable recording medium known in the art to which this invention pertains.
[0095] The structural elements of this invention, combined with a computer as hardware, can be stored as a program (or application program) on a medium. The structural elements of this invention can be programmed or executed by software elements; similarly, embodiments include data structures, processes, routines, or various algorithms implemented by combinations of other programming structures, which can be implemented using programming or scripting languages such as C, C++, Java, and assembly language. At the functional level, the algorithm can be implemented by more than one processor. Hereinafter, refer to... Figures 3 to 6 This describes the method executed by server 100 to provide matching information by analyzing sound information.
[0096] Figure 3 The flowchart illustrates a method for providing matching information by analyzing sound information according to an embodiment of the present invention.
[0097] According to one embodiment of the present invention, in step S110, the server 100 may perform the step of acquiring sound information. According to this embodiment, the sound information can be acquired through a user terminal 200 associated with the user. For example, the user terminal 200 associated with the user may include all types of handheld wireless communication devices such as smartphones, smartpads, and tablet PCs, or electronic devices (e.g., devices that receive sound information via a microphone) located in a specific space (e.g., the user's living space).
[0098] According to one embodiment of the present invention, acquiring sound information refers to receiving or loading sound information stored in a memory. Furthermore, acquiring sound information can mean receiving or loading sound information from other storage media, other servers, or additional processing modules within the same server via wired / wireless communication methods.
[0099] According to another embodiment of the present invention, the acquisition of sound information can be performed based on whether the user is located in a specific space (e.g., the user's activity space). Specifically, a sensing module can be provided in the specific space related to the user's activity. That is, the user's location in the specific space can be identified by the sensing module located in the specific space. For example, Radio Frequency Identification (RFID) technology, as a short-range communication technology, can be used to identify users at a distance by means of radio waves to share information. For example, the user can hold a card or mobile terminal including the RFID module. The RFID module held by the user records information for identifying the corresponding user (e.g., a user's personal identification identifier (ID) registered on a service management server, identification code, etc.). The sensing module can identify whether the corresponding user is located in the specific space by identifying the RFID module held by the user. In addition to RFID technology, the sensing module can include various technologies for transmitting and receiving user-specific information through contact / non-contact methods (e.g., short-range communication technologies such as Bluetooth). Furthermore, the sensing module can also include a biometric data identification module, which can identify the user's biometric data (voice, fingerprint, face) by linking with a microphone, touchpad, camera module, etc. In another embodiment, whether a user is located in a specific space can be identified through voice information related to the user's voice. Specifically, the voice information related to the user's voice can be identified as the initiation phrase, and additional sound information generated in the corresponding space can be obtained by corresponding recognition time points.
[0100] Server 100 can identify whether a user is located in a specific space through the sensing module described above or through sounds related to the user's voice. Furthermore, if it is determined that the user is located in a specific space, server 100 can acquire sound information generated at the corresponding time point.
[0101] In other words, when there are no users in a specific space, no sound information related to that space is acquired; sound information related to that space is acquired only when there are users in that specific space. This can minimize power consumption.
[0102] According to an embodiment of the present invention, the step of acquiring sound information may include the following steps: performing preprocessing on the acquired sound information; and identifying sound characteristic information corresponding to the preprocessed sound information.
[0103] According to one embodiment of the present invention, preprocessing of sound information refers to preprocessing used to improve the sound information recognition rate. For example, the preprocessing may include preprocessing to permanently remove noise from the sound information. Specifically, server 100 can perform normalization of the signal size contained in the sound information by comparing it with a standard signal size. If the signal size contained in the acquired sound information is less than a predetermined standard signal, server 100 increases the corresponding signal size; if the signal size contained in the sound information is greater than or equal to the predetermined standard signal, server 100 can perform preprocessing to decrease the corresponding signal size (i.e., to prevent clipping). The above specific description of noise removal is merely an example, and the present invention is not limited thereto.
[0104] According to another embodiment of the present invention, the preprocessing of sound information may include amplifying sounds other than vocalizations (i.e., non-verbal sounds) by analyzing the signal waveforms contained in the sound information. Specifically, the server 100 may amplify sounds associated with at least one specific frequency by analyzing multiple sound frequencies contained in the sound information.
[0105] For example, to identify the various sound types contained in the sound information, server 100 can use machine learning algorithms such as Support Vector Machine (SVM) for classification, and can amplify specific sounds using sound amplification algorithms corresponding to sounds of different frequencies. The sound amplification algorithms described above are merely examples, and the present invention is not limited thereto.
[0106] In other words, the present invention can amplify the non-verbal sounds contained in the audio information by performing preprocessing. For example, in the present invention, in order to identify user characteristics (or to provide the user with the best matching information), the analyzed audio information may include both verbal and non-verbal audio information. According to one embodiment of the present invention, at a user-specific level, non-verbal audio information can provide more meaningful analysis than verbal audio information.
[0107] As a specific example, when server 100 acquires sound information including pet (e.g., dog) sounds (i.e., non-verbal sound information), server 100 may amplify the sound information related to the corresponding non-verbal sounds in order to improve the recognition rate of "dog sounds" as non-verbal sound information.
[0108] As another example, when server 100 acquires sound information related to a user's cough sound (i.e., non-verbal sound information), server 100 may amplify the sound information related to the corresponding non-verbal sound in order to improve the recognition rate of "human cough sound" as non-verbal sound information.
[0109] In other words, by performing preprocessing to amplify nonverbal audio information to provide more meaningful information at the level of identifying user characteristics, more appropriate matching information can be provided to the user.
[0110] Furthermore, the server 100 can identify sound characteristic information corresponding to the preprocessed sound information. The sound characteristic information may include first characteristic information and second characteristic information. The first characteristic information is used to determine whether the sound information is related to at least one of speech sounds and non-speech sounds, and the second characteristic information is used to distinguish objects.
[0111] The first characteristic information may include information for determining whether the sound information belongs to speech or non-speech sounds. For example, the first characteristic information corresponding to the first sound information may include information for determining whether the corresponding first sound information is related to speech sounds, and the first characteristic information corresponding to the second sound information may include information for determining whether the second sound information is related to non-speech sounds.
[0112] The second characteristic information may include information for determining the number of objects included in the sound information. For example, the second characteristic information corresponding to the first sound information may include information that three users are making sounds in the corresponding first sound information, and the second characteristic information corresponding to the second sound information may include information that there are sounds related to washing machine operation and sounds related to cat meowing in the corresponding second sound information. In one embodiment of the present invention, the first characteristic information and the second characteristic information corresponding to the sound information can be identified by the following first sound model and second sound model.
[0113] That is, the server 100 can identify sound characteristic information corresponding to the preprocessed sound information by performing preprocessing on the acquired sound information. As mentioned above, since the sound characteristic information includes information for determining whether the corresponding sound information is related to either a speech sound or a non-speech sound (i.e., first characteristic information) and information for determining the number of objects within the sound information (i.e., second characteristic information), it provides convenience in the sound information analysis process described below.
[0114] According to one embodiment of the present invention, in step S120, the server 100 may perform a step of obtaining user characteristic information based on voice information. In one embodiment of the present invention, the step of obtaining user characteristic information may include the following steps: identifying a user or object by analyzing voice information; and generating user characteristic information based on the identified user or object.
[0115] Specifically, when acquiring sound information, server 100 can identify the user or object corresponding to the sound information by analyzing the corresponding sound information. For example, server 100 can identify the first sound information as a sound corresponding to a first user by analyzing the first sound information. As another example, server 100 can identify the second sound information as a sound corresponding to a vacuum cleaner by analyzing the second sound information. As yet another example, server 100 can identify the third sound information as containing a sound corresponding to a second user and a sound related to a washing machine by analyzing the third sound information. The specific descriptions of the first to third sound information described above are merely examples, and the present invention is not limited thereto.
[0116] Furthermore, server 100 can generate user characteristic information based on the user or object identified by the corresponding voice information. This user characteristic information is used to provide matching information; for example, it can be text related to the voice information, topic, emotion, object recognition information, or object status information.
[0117] For example, if the first sound information is identified as a sound corresponding to the first user, the server 100 can generate user characteristic information that the user is a 26-year-old female by identifying user information matching the first user. As another example, if the second sound information is identified as a sound corresponding to a vacuum cleaner of brand A, the server 100 can generate user characteristic information that the user uses a vacuum cleaner of brand A based on the corresponding second sound information. As yet another example, if the third sound information includes a sound corresponding to the second user and a sound related to a washing machine of brand B, the server 100 can generate user characteristic information that the user is a 40-year-old male who uses a washing machine of brand B. The specific descriptions of the first to third sound information and the user characteristic information corresponding to each sound information are merely examples, and the present invention is not limited thereto.
[0118] That is, server 100 can generate user characteristic information related to the user based on the user or object identified by analyzing voice information. The aforementioned user characteristic information can be information used to identify user interests, tastes, or characteristics, etc.
[0119] Furthermore, according to an embodiment of the present invention, the step of obtaining user characteristic information may include the following steps: generating activity time information related to the time a user spends in a specific space by analyzing sound information; and generating user characteristic information based on the activity time information. Specifically, the server 100 can generate activity time information related to the time a user spends in a specific space by analyzing sound information obtained corresponding to that specific space.
[0120] In one embodiment of the present invention, server 100 can obtain information for determining whether a user is located in a specific space by using sound information related to the user's voice. Specifically, server 100 identifies the speech related to the user's voice as the starting phrase, determines that the user has entered a specific space based on the identified time point, and determines that the sound information obtained in the corresponding space does not include the speech related to the user's voice. If the size of the obtained sound information is below a predetermined standard value, it can be determined that the user does not exist in the specific space. Furthermore, server 100 can generate activity time information related to the time the user is active in the specific space based on each determination time point. That is, server 100 can generate user-related activity time information by identifying whether the user is located in a specific space using sound information related to the user's voice.
[0121] In another embodiment of the present invention, server 100 can obtain information for determining whether a user is located in a specific space based on the acquired sound information magnitude. Specifically, server 100 determines that a user has entered a specific space by identifying time points where the continuously acquired sound information magnitude in the specific space exceeds a predetermined standard value, and determines that no user is present in the specific space by identifying time points where the acquired sound information magnitude in the corresponding space reaches the predetermined standard value. Furthermore, server 100 can generate activity time information related to the user's activity time in the specific space based on each determination time point. That is, server 100 identifies the sound information magnitude generated in the specific space, determines whether a user is located in the specific space based on the corresponding magnitude, and thereby generates activity time information related to the user.
[0122] In another embodiment of the present invention, server 100 can acquire information for determining whether a user is located in a specific space based on a specific activation sound. The specific activation sound can be information related to user entry and exit. For example, the activation sound can be a sound related to the opening and closing of the front door. That is, server 100 can determine whether a user is located in the corresponding space based on sound information related to unlocking the door from the outside using a password. Furthermore, server 100 can determine the presence of a user in the corresponding space based on sound information related to the opening of the internal front door. Moreover, server 100 can generate activity time information related to the time the user is active in the specific space based on each determination time point. That is, server 100 generates activity time information related to the user by identifying whether the user is located in the specific space based on activation sounds generated within the specific space.
[0123] As described above, according to various embodiments of the present invention, server 100 can generate activity time information related to the time a user is active in a specific space. For example, during a 24-hour period when a first user is active in a specific space (e.g., a residential space), server 100 can generate activity time information for the first user being active in the specific space (e.g., the residential space) for 18 hours of that 24-hour period (e.g., from 12:00 AM to 6:00 PM the following day). As another example, server 100 can generate activity time information for a second user being active in the specific space (e.g., the residential space) for 6 hours of a day (e.g., from 12:00 AM to 6:00 PM). The specific descriptions of the activity time information corresponding to the various users described above are merely examples, and the present invention is not limited thereto.
[0124] Furthermore, server 100 can generate user-specific information based on the obtained time information. For example, server 100 can generate user characteristic information indicating that the first user is relatively active in their residential space, based on the first user's activity time information during 18 hours of a 24-hour period in a specific space (e.g., a residential space). For instance, server 100 can identify that the first user spends a relatively large amount of time in their residential space based on the first user's activity time information, thereby generating user characteristic information such as the first user being a "housewife" or a "home-based worker." In another embodiment, server 100 can also more specifically deduce the user's occupation by combining the analysis information of activity time information and voice information. In this case, user characteristics can be further specified, thus making the provided matching information more accurate.
[0125] As another example, server 100 can generate user characteristic information indicating that the second user's activity in the residential space is relatively low, based on the second user's activity time information during 6 hours of a day in a specific space (e.g., a residential space). The specific descriptions of the activity time information of each user and the corresponding user characteristic information are merely examples, and the present invention is not limited thereto.
[0126] According to another embodiment of the present invention, the step of obtaining user characteristic information may include the following steps: obtaining user characteristic information based on sound characteristic information corresponding to sound information. The sound characteristic information may include first characteristic information and second characteristic information, wherein the first characteristic information is used to determine whether the sound information is related to at least one of speech sounds and non-speech sounds, and the second characteristic information is used to distinguish objects.
[0127] Specifically, the steps for obtaining user characteristic information may include at least one of the following steps: when the voice characteristic information contains first characteristic information related to speech voice, inputting voice information into a first voice model to obtain user characteristic information corresponding to the voice information; or when the voice characteristic information contains first characteristic information related to non-speech voice, inputting voice information into a second voice model to obtain user characteristic information corresponding to the voice information.
[0128] According to one embodiment of the present invention, the first sound model may be a neural network model, which learns to perform analysis on sound information related to speech sounds to identify at least one of the text, topic or emotion related to the sound information.
[0129] According to one embodiment of the present invention, the first sound model is a speech recognition model, which outputs text information corresponding to the user's voice by inputting speech information (i.e., speech sounds) contained in the sound information. It may include one or more network functions pre-learned using learning data. That is, the first sound model may include a speech recognition model that converts speech information related to the user's voice into text information. For example, the speech recognition model may input speech information related to the user's voice to output corresponding text (e.g., "There's no dog food left"). The specific descriptions of the above-mentioned speech information and its corresponding text are merely examples, and the present invention is not limited thereto.
[0130] Furthermore, the first sound model may include a text analysis model, which can use natural language processing to analyze the text information output by the corresponding speech information to grasp the context and identify the theme or emotion contained in the speech information.
[0131] In one embodiment of the present invention, the text analysis model can perform text semantic analysis on text information using a natural language processing neural network (i.e., the text analysis model) to identify keywords and grasp the theme. For example, when the text information is related to "there is no dog food left," the text analysis model can identify the theme of the corresponding sentence as "the food is all gone." The above description of the text information and its corresponding theme is merely an example, and the present invention is not limited thereto.
[0132] Furthermore, in one embodiment of the present invention, the text analysis model can output analysis values for multiple intent groups by processing text information through a natural language processing neural network. Multiple intent groups refer to sentences with specific intents distinguished according to a predetermined criterion. The natural language processing artificial neural network can use text information as input data and output intent groups to nodes by calculating the weighted values of each connection. The connection weighted values can be the weighted values of the input, output, and forget gates used in Long Short-Term Memory (LSTM) networks, or the weighted values of the general gates of Recurrent Neural Networks (RNNs). Thus, the first sound model can calculate analysis values for text information that correspond one-to-one with each intent group. Moreover, the aforementioned analysis value refers to the probability that text information can correspond to an intent group. In another embodiment, the first sound model may further include a sentiment analysis model, which performs speech analysis based on changes in voice pitch to output sentiment analysis values. That is, the first sound model may include: a speech recognition model that outputs text information corresponding to the user's speech information; a text analysis model that analyzes text information through natural language processing to grasp the topic of the sentence; and a sentiment analysis model that performs speech analysis based on changes in voice pitch to grasp the user's sentiment. Therefore, the first sound model can output text information, topic information, or emotional information related to the corresponding sound information based on the sound information contained in the speech information related to the user's speech.
[0133] Furthermore, according to embodiments of the present invention, the second sound model can be a neural network model, which learns to perform analysis on sound information related to non-verbal sounds to obtain object recognition information or object state information related to the sound information.
[0134] According to one embodiment of the present invention, the second sound model can be implemented by a learned dimensionality reduction network function and a dimensionality reduction network function learned from a complex dimensional network function, so as to output output data similar to the input data through the server 100. That is, in the structure of the learned autoencoder, the second sound model can be constructed by a dimensionality reduction network function.
[0135] According to one embodiment of the present invention, server 100 can learn an autoencoder through unsupervised learning. Specifically, server 100 can learn a dimensionality-reduced network function (e.g., an encoder) and a complex-dimensional network function (e.g., a decoder) that constitute the autoencoder to output output data similar to the input data. More specifically, by learning only the core feature data (or features) of the sound information input by the dimensionality-reduced network function during the encoding process through the hidden layer, residual information may be lost. In this case, the data output by the complex-dimensional network function through the hidden layer during the decoding process can be an approximation of the input data (i.e., the sound information), rather than a perfect copy. That is, server 100 can learn the autoencoder by adjusting the weights to make the output data as similar to the input data as possible.
[0136] An autoencoder can be a type of neural network that outputs data similar to the input data. An autoencoder may include at least one hidden layer, with an odd number of hidden layers configured between the input and output layers. The number of nodes in each layer can be symmetrically expanded from the input layer down to the intermediate layer (the bottleneck layer, or the encoding layer) to the output layer (symmetrical to the input layer). The number of input and output layers can correspond to the number of input data items remaining after preprocessing the input data. An autoencoder has a structure where the number of nodes in the encoder, including the hidden layers, gradually decreases away from the input layer. If the bottleneck layer (the layer between the encoder and decoder with the fewest nodes) has very few nodes, it may not be able to transmit a sufficient amount of information; therefore, it can also be maintained at a certain number (e.g., more than half the number of nodes in the input layer).
[0137] Server 100 can match labeled object information to store individual object feature data, using a learning dataset containing multiple learning data labeled with object information as input and output of a dimensionality reduction network for learning. Specifically, server 100 can use a dimensionality reduction network function to obtain feature data of the first object relative to the learning data contained in the first learning data subset, using a first subset of learning data labeled with first object information (e.g., a dog) as input to the dimensionality reduction network function. The obtained feature data can be represented by vectors. In this case, the feature data corresponding to the multiple learning data contained in the first learning data subset is the output based on the learning data related to the first object, and therefore, they can be located at a relatively close distance in the item space. Server 100 can store the first object information (i.e., a dog) based on the feature data related to the first object represented by vectors. In the case of the dimensionality reduction network function, the learned autoencoder can effectively extract features that allow the multidimensional network function to successfully recover the input data. Therefore, as the second voice model is implemented in the learned autoencoder through the dimensionality reduction network function, features (i.e., the voice style of each object) that can effectively recover the input data (e.g., voice information) can be extracted.
[0138] As another example, the second subset of learning data labeled with information about a second object (e.g., a cat) may contain multiple sets of learning data that can be converted into feature data and displayed in a vector space using a dimensionality reduction network function. In this case, the corresponding feature data are the outputs of the learning data related to the information about the second object (i.e., the cat), and therefore can be located at relatively close distances in the vector space. In this case, the feature data corresponding to the information about the second object can be displayed in a different vector space than the feature data corresponding to the information about the first object.
[0139] That is, when the dimensionality reduction network function that constitutes the second sound model through the above learning process takes sound information generated from a specific space (e.g., a living space) as input, it can use the dimensionality reduction network function to process the corresponding sound information to extract features corresponding to the sound information. In this case, the second sound model can evaluate the similarity of sound styles by comparing the distance between the region displaying the features corresponding to the sound information and the distance in the vector space of the feature data of each object, and can obtain the object recognition information or object state information of the corresponding sound information based on the corresponding similarity evaluation.
[0140] Specifically, the second sound model can use a dimensionality reduction network function to compute the first sound information received from the first user terminal to obtain the first feature information. In this case, the second sound model can obtain object recognition information or object state information corresponding to the first sound information based on the first feature information and by learning the positions between various object feature data pre-recorded in the vector space.
[0141] As a specific example, object recognition information related to "A brand washing machine" can be obtained based on the first object (e.g., a washing machine of brand A) that is closest to the first feature information in the vector space.
[0142] As another example, object state information related to the first sound information and the "human cough sound" can be obtained based on the space of the second object (e.g., a human cough sound) that is closest to the first feature information in the vector space. The specific descriptions of the above-mentioned object recognition information and object state information are merely examples, and the present invention is not limited thereto.
[0143] Reference Figure 4 When acquiring sound information, server 100 can determine whether the corresponding sound information is related to speech or non-speech sounds based on first characteristic information. If the sound information is speech, server 100 can apply a first sound model to identify the text, theme, or emotion corresponding to the sound information, and can obtain first user characteristic information based on the corresponding information. Furthermore, if the sound information is non-speech, server 100 applies a second sound model to obtain object recognition information or object state information corresponding to the sound information, and can obtain second user characteristic information based on the corresponding information. In other words, first user characteristic information related to the text, theme, or emotion of the corresponding speech can be obtained, and second user characteristic information related to the object recognition information or object state information of the corresponding non-speech sounds can be obtained. That is, the user characteristic information of the present invention may include first user characteristic information and second user characteristic information obtained based on whether the sound information contains speech or non-speech sounds.
[0144] According to another embodiment of the present invention, the step of obtaining user characteristic information may include the following steps: obtaining user characteristic information based on sound characteristic information corresponding to sound information. The sound characteristic information may include second characteristic information for distinguishing objects. The second characteristic information may include information for determining the number of objects contained in the sound information. For example, the second characteristic information corresponding to the first sound information may include information that three users are making sounds in the corresponding first sound information, and the second characteristic information corresponding to the second sound information may include information that there are sounds related to a washing machine operating and sounds related to a cat meowing in the corresponding second sound information.
[0145] Specifically, server 100 can obtain user characteristic information based on the second characteristic information. For example, when identifying multiple users' voices, server 100 can obtain user characteristic information about the users' lives within the characteristic space based on the second characteristic information. In this embodiment of the invention, server 100 can detect changes in the number of people in a specific space by corresponding to the second characteristic information of the sound information acquired in each time period, and can generate corresponding user characteristic information. That is, server 100 can generate user characteristic information by understanding the user's activity patterns or lifestyle patterns in a specific space through the second characteristic information.
[0146] According to an embodiment of the present invention, in step S130, the server 100 may perform the following steps: providing matching information corresponding to user characteristic information. For example, the matching information may be advertising-related information. That is, providing matching information means providing the user with advertisements that effectively increase their desire to purchase; in other words, it may mean providing the most appropriate advertising information.
[0147] As a specific example, refer to Figure 6 When the user characteristic information obtained based on specific spatially related sound information includes second user characteristic information related to the drive of the B brand washing machine 22, the server 100 can thereby provide the user 10 with matching information related to the B brand dryer. As another example, when the measured space contains second user characteristic information related to the drive of the C brand air conditioner 24, the server 100 can thereby provide the user 10 with matching information related to summer products (umbrellas or travel gear, etc.). The specific descriptions of the aforementioned user characteristic information are merely examples, and the present invention is not limited thereto.
[0148] According to an embodiment of the present invention, the step of providing matching information corresponding to user characteristic information may include the following steps: providing matching information based on the acquisition time points and frequencies of multiple user characteristic information corresponding to multiple voice information acquired within a predetermined time period. In this case, for example, the predetermined time period refers to one day (i.e., 24 hours). In other words, the server 100 may provide matching information based on the acquisition time points and frequencies of multiple user characteristic information corresponding to multiple voice information acquired based on a 24-hour period.
[0149] For example, when the same type of user characteristic information is continuously obtained at the same time (or the same time period) on a 24-hour cycle, server 100 can provide matching information at the corresponding time. As a specific example, when user characteristic information related to the drive of a brand A washing machine is obtained at the same time every day (e.g., 7 pm), server 100 can provide matching information related to a brand A dryer at 8 pm when the washing machine has finished its work.
[0150] Furthermore, for example, when a user utters a specific keyword more than a predetermined number of times (e.g., 3 times) within a 24-hour period, server 100 can provide matching information corresponding to the corresponding keyword. As a specific example, when a user utters the keyword "dog food" more than 3 times in a day, server 100 can provide matching information corresponding to dog food. The specific descriptions of the above-mentioned user characteristic information acquisition time points, frequency, and corresponding matching information are merely examples, and the present invention is not limited thereto.
[0151] In other words, server 100 can record the frequency or timing of specific keywords, thereby understanding user characteristics and providing matching information. In this case, appropriate matching information can be provided to users at the right time, thus maximizing advertising effectiveness.
[0152] According to another embodiment of the present invention, the step of providing matching information corresponding to user characteristic information may include the following steps: when the user characteristic information includes first user characteristic information and second user characteristic information, obtaining association information regarding the correlation between the first user characteristic information and the second user characteristic information; updating the matching information based on the association information; and providing the updated matching information. The first user characteristic information may be user characteristic-related information obtained based on spoken voice, and the second user characteristic information may be user characteristic-related information obtained based on non-spoken voice. That is, user characteristic information corresponding to both spoken and non-spoken voices for a given voice signal can be obtained. In this case, the server 100 may update the matching information based on the association information between the obtained user characteristic information. The association information may be information representing the correlation between spoken and non-spoken voices. Furthermore, updating the matching information can be used to further amplify the advertising effect of the matching information. For example, updating the matching information may amplify the exposure items of the matching information, or it may be an additional discount event applied to the matching information.
[0153] More specifically, the sound information may include both verbal and non-verbal sounds. In this case, a first sound model can be applied to the verbal sounds to obtain first user characteristic information related to text, topic, or emotion. Furthermore, a second sound model can be applied to the non-verbal sounds to obtain second user characteristic information related to object recognition information and object state information. Server 100 can obtain correlation information between the first and second user characteristic information. For example, the correlation information can be numerical values representing the correlation between various user characteristic information. For instance, when the first user characteristic information contains information on the topic of "dryer" and the second user characteristic information contains information related to the application of brand A dryers, server 100 can determine that the correlation between the various user characteristic information is very high and can generate correlation information corresponding to the value "98". As another example, when the first user characteristic information contains information on the topic of "dryer" and the second user characteristic information contains information related to the application of brand B washing machines, server 100 can determine that the correlation between the various user characteristic information is relatively high and can generate correlation information corresponding to the value "85". As another example, when the first user characteristic information contains information related to "vacuum cleaner" and the second user characteristic information contains information related to cat meows, the server 100 can determine that there is no correlation between the various user characteristic information pieces and can generate associated information corresponding to the value "7". The specific descriptions of the above-mentioned user characteristic information and associated information are merely examples, and the present invention is not limited thereto.
[0154] Furthermore, when the associated information exceeds a predetermined value, the server 100 can update the matching information and provide the updated matching information to the user.
[0155] As a more specific example, user speech related to verbal sounds (e.g., "Why can't it dry?", first user characteristic information) and the sound of the dryer as a non-verbal sound (i.e., second user characteristic information) can be obtained from sound information. In this case, the predetermined value is 90, and the correlation information between the various user characteristic information can be 98. In this case, the server 100 can update the matching information by recognizing that the correlation information between the various user characteristic information is above the predetermined value. That is, when highly correlated verbal and non-verbal sounds are obtained simultaneously through a single sound information, the server 100 can determine that the user has a high interest in the corresponding object and update the matching information. For example, in order to provide the user with more detailed information about the corresponding object, the server 100 can update the matching information so that the matching information includes not only dryers of brand A, but also dryers of other brands. As another example, the server 100 can update the matching information to include event information related to the discount purchase method of brand A dryers. The specific description of the above-mentioned matching information update is only an example, and the present invention is not limited thereto.
[0156] Reference Figure 5 When acquiring user characteristic information, server 100 can identify whether the corresponding user characteristic information includes first user characteristic information and second user characteristic information. In one embodiment of the present invention, when the user characteristic information acquired corresponding to voice information includes first user characteristic information, server 100 can provide matching information corresponding to the first user characteristic information. As a specific example, refer to... Figure 6 When first user characteristic information is obtained based on sound information (e.g., user vocalization), if the first user characteristic information indicates that the user has performed vocalizations related to the topic of "dog treats," then server 100 can provide matching information related to "dog treats." The specific descriptions of the aforementioned first user characteristic information and matching information are merely examples, and the present invention is not limited thereto.
[0157] Furthermore, when the user characteristic information obtained from the corresponding voice information includes second user characteristic information, the server 100 can provide matching information corresponding to the second user characteristic information. As a specific example, refer to... Figure 6 When obtaining second user characteristic information about a dog 23 located in a specific space based on sound information (e.g., a dog's bark), server 100 can provide multiple matching information related to the dog, such as dog treats, dog food, dog toys, or dog clothes. As another example, when obtaining second user characteristic information about a user being in an unhealthy state based on sound information (e.g., a user's cough), server 100 can provide multiple matching information related to the user's health, such as cold medicine, porridge, tea, or health supplements. The specific descriptions of the aforementioned second user characteristic information and matching information are merely examples, and the present invention is not limited thereto.
[0158] Furthermore, when the user characteristic information obtained for the corresponding voice information includes both the first user characteristic information and the second user characteristic information, that is, when both are included, the server 100 obtains the correlation information between the various user characteristic information pieces and can update the matching information based on the correlation information. The server 100 can also provide updated matching information.
[0159] In other words, if both related verbal and non-verbal audio are acquired simultaneously, it indicates a relatively high level of user interest. Server 100 can then provide a wealth of information for decision-making or matching information reflecting related discount events and other relevant information to increase the user's likelihood of making a purchase. That is, user interests are predicted from the acquired audio information based on the correlation between verbal and non-verbal audio, and thus, the most appropriate matching information can be provided by offering corresponding matching information. Because this method provides matching information according to user interests in an arithmetic progression, it maximizes the likelihood of a purchase conversion.
[0160] According to an embodiment of the present invention, the step of providing matching information based on user characteristic information may include the following steps: generating an environment list based on one or more user characteristic information corresponding to one or more voice information acquired according to a predetermined time period; and providing matching information based on the environment characteristic list. The environment characteristic list is information that statistically analyzes multiple user characteristic information acquired according to a predetermined time period. In one embodiment of the present invention, the predetermined time period refers to 24 hours.
[0161] In other words, server 100 uses a 24-hour time period as a benchmark to generate an environmental characteristic list related to statistical values based on user characteristic information acquired in various time periods, and can provide matching information based on the corresponding environmental characteristic list.
[0162] As a specific example, by statistically analyzing user characteristic information obtained through the environmental characteristic list at various time periods, matching information related to food or home decoration items can be provided to users who spend more time at home. As another example, for users who spend relatively more time watching television at home, server 100 can provide matching information related to newly viewed content. As yet another example, for users who spend less time at home, matching information related to health products or unmanned self-service laundry services can be provided.
[0163] Furthermore, in one embodiment of the present invention, the step of providing matching information corresponding to user characteristic information may include the following steps: identifying a first time point for providing matching information based on an environmental characteristic list; and providing matching information corresponding to the first time point. The first time point refers to the optimal time point for providing matching information to the corresponding user. That is, the optimal matching information can be provided at each time point by understanding the user's activity process in a specific space through the environmental characteristic list.
[0164] For example, when the sound of a washing machine is periodically identified in the first time period (8 p.m.) through the list of environmental characteristics, matching information related to fabric softener or dryer can be provided in the corresponding first time period.
[0165] As another example, when the sound of a vacuum cleaner is periodically identified through the environmental characteristics list during the second time period (2 p.m.), matching information related to cordless vacuum cleaners or wet cloth vacuum cleaners can be provided during the corresponding second time period.
[0166] That is, server 100 can maximize advertising effectiveness by providing appropriate matching information at specific event times. According to the various embodiments described above, server 100 can maximize advertising effectiveness by providing targeted matching information based on sound information obtained from the user's living environment.
[0167] The methods or algorithm steps described in the embodiments of this invention can be implemented directly by hardware, or by a software module executed by hardware, or by a combination thereof. The software module can reside in random access memory (RAM), read-only memory (ROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, hard disk, portable disk, CD-ROM, or any type of computer-readable recording medium known in the art to which this invention pertains.
[0168] The structural elements of this invention, combined with a computer as hardware, can be stored as a program (or application program) on a medium. The structural elements of this invention can be programmed or executed by software elements; similarly, embodiments include data structures, processes, routines, or various algorithms implemented by combinations of other programming structures, which can be implemented using programming or scripting languages such as C, C++, Java, and assembly language. At the functional level, the algorithm can be implemented by more than one processor.
[0169] While embodiments of the present invention have been described above with reference to the accompanying drawings, it should be understood that those skilled in the art can implement the present invention through other embodiments without altering the technical concept or essential features. Therefore, the embodiments described above are merely examples at all levels and should not be construed as limiting. Detailed Implementation
[0170] The above description represents the best implementation method for carrying out the present invention.
[0171] Industrial availability
[0172] This invention can be applied to the service field of providing matching information by analyzing sound information.
Claims
1. A method for providing matching information by analyzing sound information, executed in a computer device, characterized in that, Includes the following steps: Acquire sound information; Based on the aforementioned sound information, user characteristic information is obtained; and Provide matching information corresponding to the above user characteristics. The steps for obtaining sound information described above include the following: Preprocessing is performed on the acquired audio information; and Identify and perform preprocessing on the sound characteristic information corresponding to the aforementioned sound information. The aforementioned sound characteristic information includes first characteristic information and second characteristic information. The first characteristic information is used to determine whether the sound information is related to at least one of speech sounds and non-speech sounds, and the second characteristic information is used to distinguish objects. The steps for obtaining user characteristic information based on the aforementioned sound information include the following steps: obtaining the aforementioned user characteristic information based on the sound characteristic information corresponding to the aforementioned sound information. The steps to obtain the aforementioned user characteristic information include at least one of the following steps: When the aforementioned sound characteristic information includes first characteristic information related to speech sound, the aforementioned sound information is input into the first sound model to obtain user characteristic information corresponding to the aforementioned sound information; or When the aforementioned sound characteristic information includes first characteristic information related to non-verbal sounds, the aforementioned sound information is input into the second sound model to obtain user characteristic information corresponding to the aforementioned sound information.
2. The method for providing matching information by analyzing sound information according to claim 1, characterized in that, The steps for obtaining user characteristic information based on the above-mentioned sound information include the following steps: By analyzing the aforementioned audio information, users or objects can be identified; and The aforementioned user characteristic information is generated based on the identified users or objects.
3. The method for providing matching information by analyzing sound information according to claim 1, characterized in that, The steps for obtaining user characteristic information based on the above-mentioned sound information include the following steps: By analyzing the aforementioned sound information, activity time information related to the user's activity time within a specific space is generated; and The aforementioned user characteristic information is generated based on the activity time information.
4. The method for providing matching information by analyzing sound information according to claim 1, characterized in that, The steps of providing matching information corresponding to the aforementioned user characteristic information include the following steps: providing the aforementioned matching information based on the acquisition time points and frequencies of multiple user characteristic information corresponding to multiple voice information acquired within a predetermined time period.
5. The method for providing matching information by analyzing sound information according to claim 4, characterized in that, The aforementioned first sound model includes a neural network model that learns in a manner capable of identifying at least one of text, topic, or emotion related to the sound information by performing analysis on the sound information related to the speech sound. The second sound model mentioned above includes a neural network model that learns in a manner that allows it to acquire object recognition information or object state information related to the sound information by performing analysis on sound information related to non-verbal sounds. The aforementioned user characteristic information includes at least one of first user characteristic information and second user characteristic information. The first user characteristic information is at least one of text, theme, or emotion related to the aforementioned sound information, and the second user characteristic information is the aforementioned object recognition information or the aforementioned object status information related to the aforementioned sound information.
6. The method for providing matching information by analyzing sound information according to claim 5, characterized in that, The steps for providing matching information corresponding to the above user characteristic information include the following steps: When the acquired user characteristic information includes the first user characteristic information and the second user characteristic information, obtain association information regarding the correlation between the first user characteristic information and the second user characteristic information; Update the matching information based on the above-mentioned related information; as well as Provide updated matching information as described above.
7. The method for providing matching information by analyzing sound information according to claim 1, characterized in that, The steps for providing matching information based on the above user characteristic information include the following steps: Generate an environmental characteristic list based on one or more user characteristic information corresponding to one or more sound information acquired according to a predetermined time period; and The above matching information is provided based on the list of environmental characteristics mentioned above. The above list of environmental characteristics is a statistical analysis of multiple user characteristic information obtained according to the predetermined time period.
8. The method for providing matching information by analyzing sound information according to claim 7, characterized in that, The steps for providing matching information corresponding to the above user characteristic information include the following steps: Based on the aforementioned list of environmental characteristics, a first time point is identified for providing the aforementioned matching information; and The matching information mentioned above is provided corresponding to the first time point.
9. An apparatus for performing a method of providing matching information by analyzing sound information, characterized in that, include: Memory, used to store more than one instruction; and A processor for executing one or more of the instructions stored in the aforementioned memory. The processor described above executes the method according to claim 1 by executing one or more of the above instructions.
10. A computer program product stored on a computer-readable recording medium, characterized in that, Combined with a computer as hardware, it is used to perform the method according to claim 1.
Citation Information
Patent Citations
Advertisement service system capable of providing advertisement material template for auto multilink based on analyzing bigdata of user pattern and advertisement method thereof
KR102044555B1
Information service robot system
JP2004017200A
Determination device, method for determination, and determination program
JP2019036191A
Server
JP2019200598A