User profile generation methods, devices, electronic devices, and computer-readable media

By performing environmental sound event detection on user call audio, a user profile is generated, which solves the problem that existing technologies cannot determine the environment in user profiles and enables the provision of environment-related services.

CN115659211BActive Publication Date: 2026-04-07CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, user profiles generated based on user behavior data cannot determine the user's environment, resulting in the inability to provide services related to that environment.

Method used

By segmenting user call audio into audio segments without human voices and audio segments containing human voices, extracting audio features from each segment, and inputting them into a pre-trained ambient sound event detection model, a user profile is generated, which is then combined with the ambient sound event detection results.

Benefits of technology

It enriches the ways to generate user profiles, enables the determination of the user's environment, and thus provides services related to that environment, thereby improving service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115659211B_ABST
    Figure CN115659211B_ABST
Patent Text Reader

Abstract

This application discloses a user profile generation method, apparatus, electronic device, and computer-readable medium. An embodiment of the method includes: segmenting a user's call audio into an audio segment without human voice and an audio segment containing human voice; extracting a first audio feature from the audio segment without human voice, inputting the first audio feature into a pre-trained first ambient sound event detection model to obtain a first detection result; extracting a second audio feature from the audio segment containing human voice, inputting the second audio feature into a pre-trained second ambient sound event detection model to obtain a second detection result; and generating a user profile based on the first and second detection results. This implementation enriches the methods for generating user profiles. User profiles generated in this way can provide users with services related to their environment, thereby improving service quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to user profile generation methods, apparatus, electronic devices, and computer-readable media. Background Technology

[0002] User personas, also known as user profiles, are an effective tool for identifying target users and determining their needs. Through user personas, targeted services can be provided to users.

[0003] In existing technologies, user profiles are typically generated based on user behavior data within a specific application or website. However, the information provided by these user profiles is often limited, and it's impossible to determine the user's environment based on them, thus preventing the provision of services relevant to that environment. Summary of the Invention

[0004] This application proposes a user profile generation method, apparatus, electronic device, and computer-readable medium, which enriches the ways of generating user profiles. The user profiles generated in this way can provide users with services related to their environment, thereby improving service quality.

[0005] In a first aspect, embodiments of this application provide a user profile generation method, the method comprising: segmenting user call audio into an audio segment without human voice and an audio segment containing human voice; extracting a first audio feature from the audio segment without human voice, inputting the first audio feature into a pre-trained first ambient sound event detection model to obtain a first detection result; extracting a second audio feature from the audio segment containing human voice, inputting the second audio feature into a pre-trained second ambient sound event detection model to obtain a second detection result; and generating a user profile based on the first detection result and the second detection result.

[0006] Secondly, embodiments of this application provide a user profile generation device, which includes: a segmentation unit for segmenting user call audio into audio segments without human voice and audio segments containing human voice; a first detection unit for extracting a first audio feature from the audio segments without human voice and inputting the first audio feature into a pre-trained first ambient sound event detection model to obtain a first detection result; a second detection unit for extracting a second audio feature from the audio segments containing human voice and inputting the second audio feature into a pre-trained second ambient sound event detection model to obtain a second detection result; and a generation unit for generating a user profile based on the first detection result and the second detection result.

[0007] Thirdly, embodiments of this application provide an electronic device, including: one or more processors; and a storage device storing one or more programs thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in the first aspect above.

[0008] Fourthly, embodiments of this application provide a computer-readable medium having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect above.

[0009] The user profile generation method, apparatus, electronic device, and computer-readable medium provided in this application segment user call audio into audio segments without human voice and audio segments containing human voice. Then, a first audio feature is extracted from the audio segment without human voice, and the first audio feature is input into a pre-trained first environmental sound event detection model to obtain a first detection result. Next, a second audio feature is extracted from the audio segment containing human voice, and the second audio feature is input into a pre-trained second environmental sound event detection model to obtain a second detection result. Finally, a user profile is generated based on the first and second detection results. This provides a method for generating user profiles based on user call audio, enriching the methods for generating user profiles. By generating user profiles from the environmental detection results obtained by detecting the call environment of user call audio, the environment in which the user is located can be determined based on the user profile, thereby providing services related to the user's environment and improving service quality. Attached Figure Description

[0010] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0011] Figure 1 This is a flowchart of an embodiment of the user profile generation method of this application;

[0012] Figure 2 This is a schematic diagram illustrating the application scenario of the user profile generation method of this application;

[0013] Figure 3 This is a flowchart of the training process of the first and second environmental sound event detection models in this application;

[0014] Figure 4 This is a schematic diagram of the training phase of the first and second environmental sound event detection models according to this application;

[0015] Figure 5 This is a schematic diagram of the structure of an embodiment of the user profile generation device according to this application;

[0016] Figure 6 This is a schematic diagram of the structure of an electronic device used to implement the embodiments of this application. Detailed Implementation

[0017] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.

[0020] Please refer to Figure 1 The diagram illustrates a flow 100 of an embodiment of a user profile generation method according to this application. The user profile generation method includes the following steps:

[0021] Step 101: Divide the user's call audio into audio segments without human voice and audio segments containing human voice.

[0022] In this embodiment, the user profile generation method can be executed by a processor in an electronic device such as a server. The server may store user call audio.

[0023] As an example, the server mentioned above could be a carrier's server. The user's call audio can be generated during a call when the user is using services provided by the carrier. It should be noted that the carrier's server can acquire and store the user's call audio with the user's authorization.

[0024] As another example, the server mentioned above could be an instant messaging server. An instant messaging server can be used to maintain an instant messaging application. The user's call audio can be generated during a voice call conducted by the user through the instant messaging application. It should be noted that the instant messaging server can acquire and store the user's call audio with the user's authorization.

[0025] In this embodiment, the aforementioned executing entity can segment the user's call audio into silent audio segments and audio segments containing human voices. Silent audio segments can be audio segments that do not contain the user's voice and only contain background noise. Audio segments containing human voices can be audio segments that contain both the user's voice and background noise. There can be one or more silent audio segments and audio segments containing human voices. For example, if the user's call audio is one minute long, the first 5 seconds can be silent audio segments, the next 20 seconds can be audio segments containing human voices, the next 5 seconds can be silent audio segments (i.e., pauses between calls), the next 25 seconds can be audio segments containing human voices, and the last 5 seconds can be silent audio segments.

[0026] In practice, a voice separation model can be pre-trained and stored using machine learning methods (such as supervised learning). This voice separation model can be used to detect audio segments containing human voices in an audio clip. The aforementioned execution entity can input the user's call audio into the voice separation model to identify audio segments containing human voices in the user's call audio, thereby treating the remaining segments as non-human voice audio segments, thus achieving the segmentation of the user's call audio.

[0027] Step 102: Extract the first audio feature from the silent audio segment, input the first audio feature into the pre-trained first environmental sound event detection model, and obtain the first detection result.

[0028] In this embodiment, the aforementioned execution entity can first extract audio features (which can be denoted as the first audio feature) from the audio segment without voice. As an example, the audio segment without voice can be processed sequentially by audio pre-emphasis, framing, windowing, discrete Fourier transform, Mel filtering, logarithmic operation, etc., to obtain the first audio feature.

[0029] After obtaining the first audio feature, the aforementioned executing entity can input the first audio feature into a pre-trained first environmental sound event detection model to obtain a first detection result. The first environmental sound event detection model can be used to analyze the audio features of audio without human voices to output the probability that the environmental sounds in the audio belong to various environmental sound events. The first detection result may include the probability that the environmental sounds in the audio belong to various environmental sound events. The aforementioned first environmental sound event detection model can be pre-trained using machine learning methods on a multi-classification model with multi-classification capabilities. In practice, environmental sound events may include, but are not limited to, at least one of the following: car running sound events, wind noise / air conditioning sound events, background television sound events, music environmental sound events, children's sound events, etc. Each type of environmental sound event is a type of environmental sound.

[0030] Step 103: Extract the second audio features from the audio segment containing human voices, and input the second audio features into the pre-trained second environmental sound event detection model to obtain the second detection result.

[0031] In this embodiment, the aforementioned execution entity can first extract audio features (which can be denoted as the second audio feature) from the audio segment containing human voice. The method for extracting the second audio feature is basically the same as that for extracting the first audio feature, and will not be described in detail here.

[0032] After obtaining the second audio features, the aforementioned executing entity can input these features into a pre-trained second environmental sound event detection model to obtain a second detection result. This second environmental sound event detection model can be used to analyze the audio features of audio containing human voices to output the probability that the environmental sounds in the audio containing human voices belong to each environmental sound event. The second detection result can include the probability that the environmental sounds in the audio containing human voices belong to each environmental sound event. The second environmental sound event detection model can also be pre-trained using machine learning methods on a classification model with classification capabilities.

[0033] It should be noted that the second ambient sound event detection model can have the same or different structure as the first ambient sound event detection model. In practice, since audio without human voices does not contain human voice interference, it is easier to identify the type of ambient sound event compared to audio containing human voices. Therefore, the first ambient sound event detection model can be trained using a lightweight classification model (e.g., a single classification model) to reduce power consumption; the second ambient sound event detection model can be trained using a heavyweight classification model (e.g., two or more classification models) to improve the accuracy of the detection results.

[0034] In some optional implementations, the second ambient sound event detection model may include a first branch and a second branch. The first branch may include the aforementioned first ambient sound event detection model. The second branch may include a third ambient sound event detection model. The first and third ambient sound event detection models can be trained on multi-classification models with different structures. When using the second ambient sound event detection model to detect ambient sound events, the aforementioned execution entity can input the second audio features into the first and second branches respectively to obtain the first branch detection result output by the first branch and the second branch detection result output by the second branch respectively. Then, the first branch detection result and the second branch detection result can be fused to obtain the second detection result. Both the first and second branch detection results can indicate the probability that the ambient sound in the audio segment containing human voice belongs to each ambient sound event. The aforementioned execution entity can fuse the probabilities of the same ambient sound event in the first and second detection results by weighted summation, averaging, etc., to obtain the fused probability that the ambient sound in the audio segment containing human voice belongs to each ambient sound event, and use this as the second detection result.

[0035] In some alternative implementations, the first ambient sound event detection model can be trained on a discriminative network within a Generative Adversarial Network (GAN). This discriminative network can have multi-classification capabilities and can also be called a classification network. The third ambient sound event detection model can be trained on an attention-based classification model employing an attention mechanism.

[0036] Step 104: Generate a user profile based on the first and second detection results.

[0037] In this embodiment, the execution entity can fuse the first detection result and the second detection result, and based on the fused detection result, determine one or more ambient sound events to which the ambient sound in the user's call audio belongs, thereby generating a user profile based on the ambient sound event. For example, a user profile can be generated by associating the ambient sound event with a user identifier.

[0038] See Figure 2 The diagram illustrates an application scenario of the user profile generation method of this application. A user is talking to another user while riding a bus. The conversation may include ambient sounds such as bus noise and station announcements, as well as the user's voice. The aforementioned execution entity can acquire this conversation data and use it as the user's audio. First, the execution entity can segment the user's audio into two parts: a silent audio segment containing only ambient sounds such as bus noise and station announcements (when the user is not speaking or is in a speaking pause), and a voice audio segment containing both ambient sounds and the user's voice (when the user is speaking). Then, audio features can be extracted from the silent audio segment, and the extracted first audio features can be input into a pre-trained first ambient sound event detection model to obtain a first detection result. Simultaneously, audio features can be extracted from the audio segment containing the user's voice, and the extracted second audio features can be input into a pre-trained second ambient sound event detection model to obtain a second detection result. Finally, the first and second detection results can be fused to determine that the ambient sound event in the user's call audio is a car operation sound event. The label of the car operation sound event is then associated with the user identifier to generate a user profile. If a user profile already exists, the label of the car operation sound event can be associated with the user identifier to update the user profile.

[0039] In some alternative implementations, the aforementioned execution entity can generate user profiles through the following steps:

[0040] The first step is to determine the target detection result based on the first and second detection results described above. For example, the target detection result can be obtained by weighted summation or averaging of the probabilities of the same environmental sound events in the first and second detection results. The target detection result may include the target probabilities of various environmental sound events.

[0041] The second step involves determining the user's attribute information based on the target detection results and the generation time of the user's call audio. This attribute information may include, but is not limited to, occupation, industry, and mode of transportation. For example, if the user's call audio was generated during a weekday commute and the background sound is detected as car noise, the user's commute mode can be determined to be public transportation. As another example, if the user's call audio was generated during a weekday workday and the background sound is detected as wind noise or air conditioning noise, the user can be determined to be an office worker; if the background sound is detected as background television sound, the user can be determined to be a home-based worker; if the background sound is detected as ambient music, the user's occupation can be determined to be an artist; if the background sound is detected as a voice other than the user's, the user can be determined to be in the service industry; and if the background sound is detected as a child's voice, the user can be determined to be in the education industry. These will not be elaborated further here.

[0042] The second step is to generate a user profile based on the aforementioned attribute information. For example, the attribute information can be further associated with user identifiers to generate a more comprehensive user profile.

[0043] It should be noted that for the same user, this process can be executed once every time the user's call audio is processed, to detect ambient sound events, associate the tags and attribute information of the ambient sound events with the user information, thereby supplementing and updating the user profile.

[0044] Furthermore, in some optional implementations, the aforementioned entity can also determine the target time period based on the generation time period of the user's call audio. For example, the target time period could be a time period on another date that is the same as the generation time period. For instance, if the generation time period is the Monday morning commute from 7:00 to 8:00, then the target time period could be the following Monday morning from 7:00 to 8:00. Then, based on the user profile, target information can be selected and pushed to the aforementioned user. For example, the user profile may include the user's attribute information, such as a tag indicating that the commuting mode is public transportation. Suitable client applications for use during the commuting period, such as e-book applications or short video applications, can then be selected. The download method information of the aforementioned client applications can be used as target information and pushed to the user. This provides users with push services related to their environment, improving the practicality and targeting of information push.

[0045] The method provided in the above embodiments of this application involves segmenting user call audio into audio segments without human voice and audio segments containing human voice. Then, a first audio feature is extracted from the audio segment without human voice, and this first audio feature is input into a pre-trained first environmental sound event detection model to obtain a first detection result. Next, a second audio feature is extracted from the audio segment containing human voice, and this second audio feature is input into a pre-trained second environmental sound event detection model to obtain a second detection result. Finally, a user profile is generated based on the first and second detection results. This provides a method for generating user profiles based on user call audio, enriching the methods for generating user profiles. By generating user profiles from the environmental detection results obtained by detecting the call environment in user call audio, the environment in which the user is located can be determined based on the user profile, thereby providing services related to the user's environment and improving service quality.

[0046] In some alternative embodiments, see Figure 3 The flowcharts showing the training process of the first and second ambient sound event detection models are provided, along with reference to [other documentation / references]. Figure 4 The diagram illustrates the training phases of the first and second ambient sound event detection models. The first phase can be used to train the first ambient sound event detection model, which can be obtained through the following sub-steps:

[0047] Sub-step S11: Obtain the unmanned audio sample set and the random noise sample set.

[0048] Here, the silent audio sample set may include a large number of silent audio samples. Each silent audio sample does not contain the voice of the person speaking, as it includes background noise. The silent audio samples may be tagged with ambient sound events. The random noise sample set may include a large number of random noise samples, which may be random Gaussian noise, etc.

[0049] In sub-step S12, the unmanned audio samples in the unmanned audio sample set are input into the attention classification model and the discriminant network respectively. Based on the classification results output by the attention classification model and the discriminant network, the discriminant network is initially trained.

[0050] Here, the attention classification model and the discriminant network can be set up in different branches. The silent audio samples from the silent audio sample set can be input into the branch containing the attention classification model and the branch containing the discriminant network, respectively, to obtain the classification results output by the attention classification model and the discriminant network. Then, the classification results based on the environmental sound event labels, the attention classification model, and the discriminant network can be input into a preset joint loss function (which can be denoted as Loss). combineThe loss value is obtained by first calculating the loss value. Then, the parameters of the attention classification model can be fixed, and the parameters of the discriminant network can be updated based on this loss value using backpropagation and gradient descent algorithms. This step can be performed iteratively to obtain the initially trained discriminant network.

[0051] As an example, Loss combine It can be determined by the following formula:

[0052] Loss combine =α×L GD +β×L att

[0053] Where α is the weight parameter of the generative adversarial classification model's influence on the final classification result, and β is the weight parameter of the attention mechanism model's influence on the final classification result. α and β can be preset based on a large amount of data statistics. GD To generate the loss function for adversarial networks, L att Let be the loss function for the attention network.

[0054] Sub-step S13: Fix the parameters of the initially trained discriminator network, and train the generative network in the generative adversarial network based on a random noise sample set.

[0055] Here, the generative adversarial network (GAN) may include a generator network and a discriminator network after initial training. First, random noise samples from a random noise sample set are input into the generator network. Then, the generator network processes the input random noise samples and inputs the processed random noise samples into the initially trained discriminator network. The discriminator network then determines whether the input random noise samples are genuine random noise samples. After the discriminator network outputs its result, this result is input into the GAN loss function to obtain the GAN loss value. Then, the parameters of the initially trained discriminator network are fixed, and based on this loss value, the parameters of the generator network are updated using backpropagation and gradient descent algorithms. This process can be iteratively executed to obtain the trained generator network.

[0056] Sub-step S14: Fix the parameters of the trained generative network, retrain the initially trained discriminative network in the generative adversarial network using an unmanned audio sample set, and determine the retrained discriminative network as the first environmental sound event detection model.

[0057] Here, the generative adversarial network (GAN) may include a trained generative network and an initially trained discriminative network. First, silent audio samples from a silent audio sample set can be input into the trained generative network. Then, the trained generative network processes the input silent audio samples and inputs the processed samples into the initially trained discriminative network. The discriminative network then outputs a classification result. After the discriminative network outputs the classification result, this result can be input into the GAN loss function to obtain the GAN loss value. Then, the parameters of the trained generative network can be fixed, and based on this loss value, the parameters of the initially trained network are updated using backpropagation and gradient descent algorithms. This process can be iteratively executed to obtain the trained generative network.

[0058] See also Figure 3 and Figure 4 The second stage can be used to train the third environmental sound event detection model, which can be obtained through the following sub-steps:

[0059] Sub-step S15: Obtain a set of audio samples containing human voices.

[0060] Here, the audio sample set containing human voices can contain a large number of audio samples containing human voices. Each audio sample containing human voices can include both ambient sound and human voices. The audio samples containing human voices can be collected in real-world scenarios and can be compiled based on known sample sets. The audio samples containing human voices can be tagged with ambient sound events.

[0061] Optionally, the audio samples containing human voices in the audio sample set are generated through the following steps: First, extract the audio samples without human voices from the audio sample set without human voices, and extract the pure human voice samples from the audio sample set without human voices; then, synthesize the extracted audio samples without human voices and the pure human voice samples to generate the audio samples containing human voices.

[0062] In sub-step S16, the audio samples containing human voices in the set of audio samples containing human voices are input into the attention classification model and the first environmental sound event detection model, respectively. Based on the classification results output by the attention classification model and the first environmental sound event detection model, the attention classification model is trained to obtain the third environmental sound event detection model.

[0063] Here, the attention classification model and the first ambient sound event detection model can be set up in different branches. Audio samples containing human voices from the set of audio samples are input into the branches containing the attention classification model and the first ambient sound event detection model, respectively, to obtain the classification results output by the attention classification model and the first ambient sound event detection model. Then, the ambient sound event labels attached to the audio samples containing human voices, the classification results output by the attention classification model, and the classification results output by the first ambient sound event detection model can be input into a preset joint loss function (i.e., Loss). combine The loss value is obtained. Then, the parameters of the first ambient sound event detection model can be fixed, and based on this loss value, the parameters of the attention classification model are updated using backpropagation and gradient descent algorithms. This step can be iteratively executed to obtain the trained attention classification model, which is then used as the second ambient sound event detection model.

[0064] By using a single first ambient sound event detection model to detect ambient sound events in audio segments without human voices, and by using a dual model that includes both the first and third ambient sound event detection models to detect ambient sound events in audio segments containing human voices, the detection power consumption for ambient sound events in audio segments without human voices can be reduced while the accuracy of ambient sound event detection in audio segments containing human voices can be improved.

[0065] Further reference Figure 5 As an implementation of the methods shown in the above figures, this application provides an embodiment of a user profile generation device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0066] like Figure 5 As shown, the user profile generation device 500 of this embodiment includes: a segmentation unit 501, used to segment user call audio into audio segments without human voice and audio segments containing human voice; a first detection unit 502, used to extract a first audio feature from the audio segments without human voice, and input the first audio feature into a pre-trained first ambient sound event detection model to obtain a first detection result; a second detection unit 503, used to extract a second audio feature from the audio segments containing human voice, and input the second audio feature into a pre-trained second ambient sound event detection model to obtain a second detection result; and a generation unit 504, used to generate a user profile based on the first detection result and the second detection result.

[0067] In some optional implementations, the second ambient sound event detection model includes a first branch and a second branch, the first branch including the first ambient sound event detection model, and the second branch including a third ambient sound event detection model; the second detection unit 503 is further configured to input the second audio feature to the first branch and the second branch to obtain the first branch detection result output by the first branch and the second branch detection result output by the second branch; and to fuse the first branch detection result and the second branch detection result to obtain the second detection result.

[0068] In some alternative implementations, the first ambient sound event detection model is obtained by training a discriminative network in a generative adversarial network; the third ambient sound event detection model is obtained by training an attention classification model that employs an attention mechanism.

[0069] In some optional implementations, the first ambient sound event detection model is trained through the following steps: acquiring a set of unmanned audio samples and a set of random noise samples; inputting the unmanned audio samples from the unmanned audio sample set into the attention classification model and the discriminant network respectively, and performing initial training on the discriminant network based on the classification results output by the attention classification model and the discriminant network; fixing the parameters of the discriminant network after initial training, and training the generative network in the generative adversarial network based on the random noise sample set; fixing the parameters of the trained generative network, and retraining the initially trained discriminant network in the generative adversarial network using the unmanned audio sample set, and determining the retrained discriminant network as the first ambient sound event detection model.

[0070] In some optional implementations, the aforementioned third ambient sound event detection model is trained through the following steps: obtaining a set of audio samples containing human voices; inputting the audio samples containing human voices from the aforementioned set of audio samples containing human voices into the aforementioned attention classification model and the aforementioned first ambient sound event detection model respectively; and training the aforementioned attention classification model based on the classification results output by the aforementioned attention classification model and the aforementioned first ambient sound event detection model to obtain the third ambient sound event detection model.

[0071] In some optional implementations, the audio samples containing human voices in the above-mentioned audio sample set are generated by the following steps: extracting unvoiceable audio samples from the above-mentioned unvoiceable audio sample set, and extracting pure human voice samples from the above-mentioned pure human voice audio sample set; and synthesizing the extracted unvoiceable audio samples and pure human voice samples to generate audio samples containing human voices.

[0072] In some optional implementations, the generation unit 504 is further configured to determine the target detection result based on the first detection result and the second detection result; determine the user's attribute information based on the target detection result and the generation time period of the user's call audio; and generate a user profile based on the attribute information.

[0073] In some optional implementations, the above-mentioned device further includes a push unit, which is used to determine a target time period based on the generation time period of the user's call audio; select target information based on the user profile; and push the target information to the user during the target time period.

[0074] The apparatus provided in the above embodiments of this application divides user call audio into audio segments without human voice and audio segments containing human voice. Then, it extracts a first audio feature from the audio segment without human voice and inputs the first audio feature into a pre-trained first environmental sound event detection model to obtain a first detection result. Next, it extracts a second audio feature from the audio segment containing human voice and inputs the second audio feature into a pre-trained second environmental sound event detection model to obtain a second detection result. Finally, it generates a user profile based on the first and second detection results. This provides a method for generating user profiles based on user call audio, enriching the methods for generating user profiles. By generating user profiles from the environmental detection results obtained by detecting the call environment of user call audio, it is possible to determine the user's environment based on the user profile, thereby providing services related to the user's environment and improving service quality.

[0075] The following is for reference. Figure 6 It shows a schematic diagram of the structure of an electronic device used to implement some embodiments of this application. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this application.

[0076] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0077] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, disks, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.

[0078] In particular, according to some embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from ROM 602. When the computer program is executed by processing device 601, it performs the functions defined above in the methods of some embodiments of this application.

[0079] It should be noted that the computer-readable medium described in some embodiments of this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0080] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0081] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: segment user call audio into audio segments without human voice and audio segments containing human voice; extract a first audio feature from the audio segments without human voice, input the first audio feature into a pre-trained first ambient sound event detection model, and obtain a first detection result; extract a second audio feature from the audio segments containing human voice, input the second audio feature into a pre-trained second ambient sound event detection model, and obtain a second detection result; and generate a user profile based on the first and second detection results.

[0082] Computer program code for performing operations of some embodiments of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++; and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, or it can be connected to an external computer (e.g., via the Internet using an Internet service provider), including local area networks (LANs) or wide area networks (WANs).

[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0084] The units described in some embodiments of this application can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a first determining unit, a second determining unit, a selecting unit, and a third determining unit. The names of these units do not necessarily limit the specific unit itself.

[0085] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0086] The above description is merely a selection of preferred embodiments of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this application.

Claims

1. A method for generating user profiles, characterized in that, The method includes: The user's call audio is segmented into audio segments without human voice and audio segments containing human voice. Extract a first audio feature from the unmanned audio segment, and input the first audio feature into a pre-trained first environmental sound event detection model to obtain a first detection result; Extract the second audio feature from the audio segment containing human voice, and input the second audio feature into the pre-trained second environmental sound event detection model to obtain the second detection result; Based on the first detection result and the second detection result, a user profile is generated; The second ambient sound event detection model includes a first branch and a second branch, and the first branch includes the first ambient sound event detection model. The second branch includes a third environmental event detection model; The first environmental sound event detection model and the third environmental event detection model are obtained by training multi-classification models with different structures. The second audio feature is input into the first branch and the second branch respectively to obtain the first branch detection result output by the first branch and the second branch detection result output by the second branch respectively. The first branch detection result and the second branch detection result are fused to obtain the second detection result. The first ambient sound event detection model is obtained by training the discriminant network in a generative adversarial network, and the third ambient sound event detection model is obtained by training an attention classification model using an attention mechanism. The step of generating a user profile based on the first detection result and the second detection result includes: Based on the first detection result and the second detection result, the target detection result is determined; Based on the target detection results and the generation time period of the user's call audio, the user's attribute information is determined; Based on the attribute information, a user profile is generated.

2. The method according to claim 1, characterized in that, The first ambient sound event detection model is trained through the following steps: Obtain a set of audio samples with no human voices and a set of random noise samples; The unmanned audio samples in the unmanned audio sample set are respectively input into the attention classification model and the discriminant network. Based on the classification results output by the attention classification model and the classification results output by the discriminant network, the discriminant network is initially trained. With the parameters of the initially trained discriminative network fixed, the generative network in the generative adversarial network is trained based on the random noise sample set; The parameters of the generator network after fixed training are used to retrain the discriminator network after initial training in the generative adversarial network using the unmanned voice audio sample set. The retrained discriminator network is then determined as the first environmental sound event detection model.

3. The method according to claim 2, characterized in that, The third ambient sound event detection model is trained through the following steps: Obtain a set of audio samples containing human voices; The audio samples containing human voices in the set of audio samples containing human voices are respectively input into the attention classification model and the first environmental sound event detection model. Based on the classification results output by the attention classification model and the first environmental sound event detection model, the attention classification model is trained to obtain the third environmental sound event detection model.

4. The method according to claim 3, characterized in that, The audio samples containing human voices in the set of audio samples containing human voices are generated through the following steps: Extract unmanned audio samples from the unmanned audio sample set, and extract pure human voice samples from the pure human voice audio sample set; The extracted audio samples without human voices and the pure human voice samples are combined to generate audio samples containing human voices.

5. The method according to claim 1, characterized in that, The method further includes: Based on the generation time period of the user's call audio, determine the target time period; Based on the user profile, target information is selected, and the target information is pushed to the user during the target time period.

6. A user profile generation device, characterized in that, The device includes: The segmentation unit is used to divide the user's call audio into audio segments without human voice and audio segments containing human voice. The first detection unit is used to extract a first audio feature from the unmanned audio segment, input the first audio feature into a pre-trained first environmental sound event detection model, and obtain a first detection result. The second detection unit is used to extract a second audio feature from the audio segment containing human voice, input the second audio feature into a pre-trained second environmental sound event detection model, and obtain a second detection result. The generation unit is used to generate a user profile based on the first detection result and the second detection result; The device is also used to include a first branch and a second branch in the second ambient sound event detection model, wherein the first branch includes the first ambient sound event detection model; The second branch includes a third environmental event detection model; The first environmental sound event detection model and the third environmental event detection model are obtained by training multi-classification models with different structures. The second audio feature is input into the first branch and the second branch respectively to obtain the first branch detection result output by the first branch and the second branch detection result output by the second branch respectively. The first branch detection result and the second branch detection result are fused to obtain the second detection result. The first ambient sound event detection model is obtained by training the discriminant network in a generative adversarial network, and the third ambient sound event detection model is obtained by training an attention classification model that employs an attention mechanism. The device is further configured to determine a target detection result based on the first detection result and the second detection result; Based on the target detection results and the generation time period of the user's call audio, the user's attribute information is determined; Based on the attribute information, a user profile is generated.

7. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Speaker classification method and device in video, electronic equipment and storage medium

    CN113343831A

  • Audio system for artificial reality applications

    US20220182772A1