A method and apparatus for identifying user attributes

By performing convolution feature extraction and timing feature analysis on the voice fragments of the target user, combined with probability fusion, efficient and accurate user attribute recognition is achieved, solving the problems of large amount of computing and privacy leakage in the prior art.

CN113362852BActive Publication Date: 2025-07-11SHENZHEN WANGYU COMPUTER NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010142092.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-04
Publication Date
2025-07-11
Estimated Expiration
2040-03-04

AI Technical Summary

Technical Problem

When the prior art recognizes user attributes by collecting user images, privacy is easily leaked and the calculation amount is large, making it difficult to efficiently and accurately identify user attribute categories.

Method used

Multiple voice fragments with time-series relationships of the target user are obtained, and feature extraction and fusion are performed through convolutional neural networks and recurrent neural networks, the attribute category probability of the voice fragments is predicted, and the attribute category of the target user is determined based on probability fusion.

Benefits of technology

It improves the computing speed and result accuracy of user attribute recognition, reduces the amount of calculation, prevents privacy leakage, and improves information security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113362852B_ABST
    Figure CN113362852B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for user attribute recognition; the present application can obtain multiple speech segments of a target user with a temporal relationship, perform convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment, extract temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category based on the temporal feature information of each speech segment, fuse the probabilities of the predicted speech segments based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category, and determine the target attribute category corresponding to the target user from the preset attribute categories based on the fused probability corresponding to each preset attribute category; by improving the method for identifying the attribute category of the target user, the present application can improve the operation speed and result accuracy of user attribute recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and device for identifying user attributes. Background Art

[0002] With the rapid development of information technology, the development of network media has become more diverse and convenient. Users with different attributes have different demands for network media. In some cases, it is necessary to obtain user attributes, such as the age and gender of users, in order to provide corresponding functions or content according to user attributes. For example, for online games, since online games may affect the mental health development of underage users, it is necessary to restrict the use of online games by underage users. Another example is that for the scenario of advertising placement, corresponding advertising content can be recommended according to the gender of users.

[0003] In the current related technologies, user attributes are generally identified by collecting user images. However, images often include some relatively private information of users. Identifying user attributes by collecting user images is likely to disclose user privacy and have a negative impact on users. Summary of the Invention

[0004] Embodiments of this application provide a method and device for identifying user attributes, which can improve the operation speed and result accuracy of user attribute identification.

[0005] Embodiments of this application provide a method for identifying user attributes, including:

[0006] Obtaining multiple speech segments of a target user with a temporal relationship;

[0007] Performing convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment;

[0008] Based on the two-dimensional convolutional speech features of each speech segment, extracting temporal feature information of each speech segment;

[0009] Based on the temporal feature information of each speech segment, predicting the probability that the attribute category of each speech segment is each preset attribute category;

[0010] Based on the preset attribute categories, fusing the predicted probabilities of each speech segment to obtain fused probabilities corresponding to each preset attribute category;

[0011] Based on the fused probabilities corresponding to each preset attribute category, determining the target attribute category corresponding to the target user from the preset attribute categories.

[0012] Correspondingly, embodiments of this application provide a device for identifying user attributes, including:

[0013] An acquisition unit, configured to acquire multiple speech segments with a temporal relationship of a target user;

[0014] A first extraction unit, configured to perform convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment;

[0015] A second extraction unit, configured to extract temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment;

[0016] A prediction unit, configured to predict the probability that the attribute category of each speech segment is each preset attribute category based on the temporal feature information of each speech segment;

[0017] A fusion unit, configured to fuse the predicted probabilities of each speech segment based on the preset attribute category to obtain a fused probability corresponding to each preset attribute category;

[0018] A determination unit, configured to determine a target attribute category corresponding to the target user from the preset attribute categories based on the fused probabilities corresponding to each preset attribute category.

[0019] Optionally, in some embodiments of the present application, the first extraction unit may include a division sub-unit, a transformation sub-unit, and a first extraction sub-unit, as follows:

[0020] The division sub-unit is configured to divide each speech segment into at least one speech frame;

[0021] The transformation sub-unit is configured to transform the speech frame from the time domain to the frequency domain to obtain sub-spectrum information of the speech frame;

[0022] The first extraction sub-unit is configured to perform convolutional feature extraction on the sub-spectrum information of the speech frames of each speech segment to obtain two-dimensional convolutional speech features of each speech segment.

[0023] Optionally, in some embodiments of the present application, the first extraction sub-unit may specifically be configured to fuse the sub-spectrum information of each speech frame in each speech segment to obtain spectrum information of each speech segment; perform convolutional feature extraction on the spectrum information of each speech segment to obtain two-dimensional convolutional speech features of each speech segment.

[0024] Optionally, in some embodiments of the present application, the first extraction unit may include a convolutional sub-unit and a dimensionality reduction sub-unit, as follows:

[0025] The convolutional sub-unit is configured to perform convolutional feature extraction on each speech segment to obtain a two-dimensional convolutional feature map of each speech segment;

[0026] A dimensionality reduction sub-unit, configured to perform a dimensionality reduction operation on the two-dimensional convolutional feature maps of each speech segment to obtain two-dimensional convolutional speech features of each speech segment.

[0027] Optionally, in some embodiments of the present application, the second extraction unit may include an acquisition sub-unit, a second extraction sub-unit, an update sub-unit, and a return sub-unit, as follows:

[0028] The acquisition sub-unit is configured to acquire historical timing features of historical speech segments;

[0029] The second extraction sub-unit is configured to extract timing feature information of the current speech segment based on the two-dimensional convolutional speech features of the current speech segment and the historical timing features of the historical speech segments;

[0030] The update sub-unit is configured to update the historical timing features of the historical speech segments based on the two-dimensional convolutional speech features of the current speech segment;

[0031] The return sub-unit is configured to use the next speech segment as the new current speech segment, and return to execute the step of acquiring the historical timing features of the historical speech segments until the timing feature information of each speech segment is obtained.

[0032] Optionally, in some embodiments of the present application, the first extraction unit may specifically perform convolutional feature extraction on each speech segment through a two-dimensional convolutional neural network to obtain two-dimensional convolutional speech features of each speech segment; the second extraction unit may specifically use a recurrent neural network to extract the timing feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment.

[0033] Optionally, in some embodiments of the present application, the user attribute recognition device may further include a training unit, as follows:

[0034] The training unit is specifically configured to obtain training data, where the training data includes sample speech segments and the target attribute categories corresponding to the sample speech segments; perform convolutional feature extraction on the sample speech segments through a two-dimensional convolutional neural network to obtain two-dimensional convolutional speech features of the sample speech segments; use a recurrent neural network to extract the timing feature information of the sample speech segments based on the two-dimensional convolutional speech features of the sample speech segments; predict the probabilities that the attribute category of the sample speech segment is each preset attribute category based on the timing feature information of the sample speech segment; adjust the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probability that the predicted attribute category of the sample speech segment is the target attribute category meets a preset condition.

[0035] Optionally, in some embodiments, the step of "adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probability that the predicted attribute category of the sample speech segment is the target attribute category meets a preset condition" may specifically include:

[0036] Calculating the loss value between the probability that the predicted attribute category of the sample speech segment is each preset attribute category and the true probability, where the true probability that the attribute category of the sample speech segment is the target attribute category is 1, and the true probability that the attribute category of the sample speech segment is other preset attribute categories except the target attribute category is 0;

[0037] Based on the loss value, adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the calculated loss value is less than a preset loss value.

[0038] Optionally, in some embodiments of the present application, the attribute category includes age group and gender; the prediction unit may specifically be used to predict the probability that the age group corresponding to each speech segment is each preset age group and the probability that the corresponding gender is each preset gender based on the temporal feature information of each speech segment.

[0039] Optionally, in some embodiments of the present application, the attribute category includes age group and gender; the fusion unit may specifically be used to fuse the probabilities of the age groups corresponding to the predicted speech segments based on the preset age groups to obtain the fused probabilities corresponding to each preset age group; and fuse the probabilities of the genders corresponding to the predicted speech segments based on the preset genders to obtain the fused probabilities corresponding to each preset gender.

[0040] Optionally, in some embodiments of the present application, the attribute category includes age group and gender; the determination unit may specifically be used to determine the target age group corresponding to the target user from the preset age groups based on the fused probabilities corresponding to each preset age group; and determine the target gender corresponding to the target user from the preset genders based on the fused probabilities corresponding to each preset gender.

[0041] An electronic device provided by an embodiment of the present application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the user attribute recognition method provided by the embodiment of the present application.

[0042] In addition, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the user attribute recognition method provided by the embodiment of the present application are implemented.

[0043] The embodiments of the present application provide a method and device for user attribute recognition, which can obtain multiple speech segments of a target user with a temporal relationship, extract convolutional features from each speech segment to obtain two-dimensional convolutional speech features of each speech segment, extract temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category based on the temporal feature information of each speech segment, fuse the probabilities of the predicted speech segments based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category, and determine the target attribute category corresponding to the target user from the preset attribute categories based on the fused probability corresponding to each preset attribute category; by improving the method for recognizing the attribute category of the target user, the present application can improve the operation speed and result accuracy of user attribute recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1a is a schematic diagram of the scenario of the user attribute recognition method provided by the embodiments of the present application;

[0046] Figure 1b is a flowchart of the user attribute recognition method provided by the embodiments of the present application;

[0047] Figure 1c is a schematic diagram of the model of the user attribute recognition method provided by the embodiments of the present application;

[0048] Figure 1d is another flowchart of the user attribute recognition method provided by the embodiments of the present application;

[0049] Figure 1e is a model training diagram of the user attribute recognition method provided by the embodiments of the present application;

[0050] Figure 2a is another flowchart of the user attribute recognition method provided by the embodiments of the present application;

[0051] Figure 2b is another flowchart of the user attribute recognition method provided by the embodiments of the present application;

[0052] Figure 2c is a schematic diagram of the interface of the user attribute recognition method provided by the embodiments of the present application;

[0053] Figure 2d It is another schematic diagram of the user attribute recognition method provided by the embodiment of the present application;

[0054] Figure 3a It is a schematic structural diagram of the user attribute recognition device provided by the embodiment of the present application;

[0055] Figure 3b It is another schematic structural diagram of the user attribute recognition device provided by the embodiment of the present application;

[0056] Figure 3c It is another schematic structural diagram of the user attribute recognition device provided by the embodiment of the present application;

[0057] Figure 3d It is another schematic structural diagram of the user attribute recognition device provided by the embodiment of the present application;

[0058] Figure 3e It is another schematic structural diagram of the user attribute recognition device provided by the embodiment of the present application;

[0059] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application;

[0060] Figure 5 It is an optional schematic structural diagram of the distributed system 100 applied to the blockchain system provided by the embodiment of the present application;

[0061] Figure 6 It is an optional schematic diagram of the block structure provided by the embodiment of the present application. Detailed implementation manners

[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0063] The embodiment of the present application provides a user attribute recognition method and device. Specifically, the embodiment of the present application provides a user attribute recognition device applicable to an electronic device, and the electronic device can be a device such as a terminal or a server.

[0064] It can be understood that the user attribute recognition method in this embodiment can be executed on a terminal, or on a server, or jointly executed by a terminal and a server.

[0065] Reference Figure 1a, taking the example of the user attribute recognition method jointly executed by the terminal and the server. The user attribute recognition system provided by the embodiments of this application includes the terminal 10 and the server 11, etc.; the terminal 10 and the server 11 are connected through a network, for example, through a wired or wireless network connection, etc. Among them, the user attribute recognition device can be integrated in the server.

[0066] Among them, the terminal 10 can obtain multiple speech segments of the target user with a time sequence relationship through the voice input module, and send the speech segments to the server 11, so that the server 11 can process and analyze the speech segments based on the multiple speech segments of the target user received with a time sequence relationship, obtain the target attribute category corresponding to the target user, and then return the target attribute category corresponding to the target user to the terminal 10. Among them, the terminal 10 can include a mobile phone, a smart TV, a tablet computer, a laptop computer, or a personal computer (PC, Personal Computer), etc.

[0067] The server 11 can be used to: obtain multiple speech segments of the target user with a time sequence relationship, extract convolutional features for each speech segment to obtain two-dimensional convolutional speech features of each speech segment, extract time sequence feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category based on the time sequence feature information of each speech segment, fuse the probabilities of the predicted speech segments based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category, determine the target attribute category corresponding to the target user from the preset attribute categories based on the fused probability corresponding to each preset attribute category, and then send the determined target attribute category corresponding to the target user to the terminal 10. Among them, the server 11 can be a single server or a server cluster composed of multiple servers.

[0068] The process of the above server 11 determining the target attribute category corresponding to the target user can also be executed by the terminal 10.

[0069] The user attribute recognition method provided by the embodiments of this application involves speech technology (Speech Technology) and machine learning (ML, Machinelearning) in the field of artificial intelligence (AI, Artificial Intellegence). The embodiments of this application can improve the operation speed and result accuracy of user attribute recognition by improving the method of recognizing the attribute category of the target user.

[0070] Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technology and software-level technology. Among them, artificial intelligence software technology mainly includes computer vision technology, speech technology, natural language processing technology, and machine learning / deep learning and other directions.

[0071] Among them, the key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction. Among them, speech has become one of the most promising human-computer interaction methods in the future.

[0072] Machine learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0073] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0074] Embodiment 1:

[0075] The embodiment of the present application will be described from the perspective of the user attribute recognition device. This user attribute recognition device can be specifically integrated in an electronic device, which can be a server or a terminal device, etc.

[0076] The user attribute recognition method according to the embodiments of the present application can be applied to various scenarios that require user attribute recognition. For example, it can be applied to the scenario of advertising placement. Through the user attribute recognition method provided in this embodiment, the age, gender, etc. of the user can be recognized, and corresponding advertising content can be recommended according to the age and gender of the user, which can improve the accuracy of advertising placement. For another example, it can be applied to online games. Young users are prone to being addicted to online games, which has an impact on the mental health development. Therefore, it is necessary to restrict the use of online games by young users. The user attribute recognition method provided in this embodiment can be used to identify young users, and then the daily login time of young users to online games can be restricted. For another example, it can also be applied to auxiliary tools for reconnaissance and evidence analysis.

[0077] As Figure 1b shown, the specific process of this user attribute recognition method is as follows. This user attribute recognition method can be executed by a server or by a terminal, and this embodiment does not make any restrictions on this.

[0078] 101. Obtain multiple speech segments of a target user with a temporal relationship.

[0079] In this embodiment, the target user is the user for whom the attribute category needs to be obtained. Among them, the attribute category can be age, gender, etc., and this embodiment does not make any restrictions on this.

[0080] Optionally, the obtained speech segments of the target user can be preprocessed. That is, before the step of "obtaining multiple speech segments of a target user with a temporal relationship", it may include:

[0081] Obtain the audio information of the target user, where the audio information includes multiple speech segments of the target user with a temporal relationship;

[0082] Perform denoising processing on the audio information to obtain multiple speech segments of the target user with a temporal relationship.

[0083] Among them, the audio information of the target user collected can be stored in the pulse code modulation (PCM) digital format as s. Optionally, multiple speech segments s containing the discrete speech of the target user can be captured from the audio information s through a voice activity detector (VAD). v Each speech segment s v can be labeled according to the temporal relationship as That is and and so on, j ∈ {1, 2, 3... J}, that is, j can be natural numbers 1, 2, 3, 4, etc., representing the jth speech segment, t jRepresents different time period labels, and each voice segment s v Can be T seconds, and T can be set according to actual needs. This embodiment does not limit this, for example, it can be set to 3 seconds. Among them, voice activity detection is to detect whether there is a voice signal in the current audio information, that is, to judge the input signal, and different processing methods are adopted for the two signals of the voice signal and various background noise signals respectively to distinguish the voice signal from various background noise signals. Distinguishing the voice signal and the background noise signal through VAD is beneficial to improving the anti-noise ability of the user attribute recognition device and better overcoming the interference of environmental noise.

[0084] 102. Perform convolutional feature extraction on each voice segment to obtain the two-dimensional convolutional voice features of each voice segment.

[0085] In this embodiment, the step "Perform convolutional feature extraction on each voice segment to obtain the two-dimensional convolutional voice features of each voice segment" may include:

[0086] Divide each voice segment into at least one voice frame;

[0087] Transform the voice frame from the time domain to the frequency domain to obtain the sub-spectrum information of the voice frame;

[0088] Perform convolutional feature extraction on the sub-spectrum information of the voice frames of each voice segment to obtain the two-dimensional convolutional voice features of each voice segment.

[0089] Among them, to transform the voice frame from the time domain to the frequency domain, the short-time discrete Fourier transform (STDFT, Short Time Discrete Fourier Transform) technology or the short-time discrete cosine transform technology (STDCT, Short Time Discrete Cosine Transform) can be used to perform time-frequency orthogonal decomposition on the voice frame signal to obtain the sub-spectrum information of the voice frame.

[0090] Optionally, each voice segment s v Is T seconds, and the T-second voice segment s v Can be divided into frames with k milliseconds as one frame. The value of k can be set according to the actual situation. For example, it can be set to 10 - 20 milliseconds. The T-second voice segment is divided into N k-millisecond voice frames, where the value of N is as shown in formula (1):

[0091]

[0092] Indicates taking The integer part of the value.

[0093] Optionally, in some embodiments, the step of "performing convolutional feature extraction on the sub-spectrum information of the speech frames of each speech segment to obtain the two-dimensional convolutional speech features of each speech segment" may include:

[0094] Fusing the sub-spectrum information of each speech frame in each speech segment to obtain the spectrum information of each speech segment;

[0095] Performing convolutional feature extraction on the spectrum information of each speech segment to obtain the two-dimensional convolutional speech features of each speech segment.

[0096] Among them, fusing the sub-spectrum information of each speech frame in each speech segment may specifically be splicing the sub-spectrum information of each speech frame in each speech segment to obtain the spectrum information of each speech segment. For example, splicing can be performed according to the time sequence relationship of the speech frames.

[0097] Optionally, by dividing a T-second speech segment, each speech segment correspondingly obtains N speech frames of k milliseconds. Transforming the N speech frames of each speech segment from the time domain to the frequency domain, that is, orthogonally decomposing and projecting the N speech frames of each speech segment in the time-frequency domain onto N subbands, and each subband is s n,m , which is the sub-spectrum information corresponding to each speech frame. Among them, n represents the serial number of the speech frame. For example, s 1,m represents the first speech frame, s 2,m represents the second speech frame, s n,m represents the Nth speech frame; m can represent the size of each subband. Then, fusing the sub-spectrum information of each speech frame in each speech segment may specifically splice the sub-spectrum information s n,m of each speech frame in each speech segment to obtain a time-frequency speech spectrogram matrix, that is, the signals of each subband in the time-frequency domain within T seconds can be represented by the time-frequency speech spectrogram matrix, as shown in formula (2):

[0098] S N×M =[|s n,m |] (2)

[0099] Among them, |*| is the modulus operation, and S N×M indicates that this time-frequency speech spectrogram matrix is N rows and M columns. Then, feature extraction can be performed on the fused S N×M .

[0100] Optionally, in some embodiments, the step of "performing convolutional feature extraction on each speech segment to obtain the two-dimensional convolutional speech features of each speech segment" may include:

[0101] Performing convolutional feature extraction on each speech segment to obtain the two-dimensional convolutional feature map of each speech segment;

[0102] Perform a dimensionality reduction operation on the two-dimensional convolutional feature maps of each speech segment to obtain the two-dimensional convolutional speech features of each speech segment.

[0103] Among them, convolutional feature extraction can be performed on each speech segment through a neural network, and this neural network can be a convolutional neural network (CNN, Convolutional Neural Networks), a visual geometry group network (VGGNet, Visual Geometry Group Network), a residual network (ResNet, Residual Network), a dense connection convolutional network (DenseNet, Dense Convolutional Network), etc. However, it should be understood that the neural network in this embodiment is not limited to the several types listed above. In addition, other machine learning methods can also be used to obtain the feature information of speech segments, such as various acoustic features like Mel spectrogram and pitch, not limited to the method of artificial neural network.

[0104] Optionally, dimensionality reduction can be performed on the two-dimensional convolutional feature maps of each speech segment by performing pooling on the two-dimensional convolutional feature maps of each speech segment. Among them, this pooling can include max-pooling (Max-pooling, Maximum Pooling), average pooling (Avg-pooling, Average Pooling), generalized mean pooling (GEM-pooling, Generalized-mean Pooling), etc.

[0105] For example, in some embodiments, the step of "performing convolutional feature extraction on each speech segment to obtain the two-dimensional convolutional speech features of each speech segment" may include:

[0106] Performing convolutional feature extraction on each speech segment through a two-dimensional convolutional neural network to obtain the two-dimensional convolutional speech features of each speech segment.

[0107] Among them, a two-dimensional convolutional neural network can be used as a feature extractor to perform feature extraction on the S after sub-spectrum information fusion N×M to obtain the two-dimensional convolutional speech features of each speech segment. If the sampling rate of the audio information is 16000Hz and the length of the frequency reference value is 48 to 96 frequency points, then each speech frame is about 3 to 6 milliseconds. Optionally, the number of convolutional filters of the two-dimensional convolutional neural network can be 16 to 32, the size of each convolutional filter is 4*4, and appropriate pooling sizes can be selected for each layer of the convolutional layer according to actual needs to perform dimensionality reduction on the two-dimensional convolutional feature maps of speech segments. For example, a pooling layer with a size of 2*4 can be selected. Finally, the output of the last layer of the convolutional layer is used as the feature output D I,J, D I,J That is, the two-dimensional convolutional speech feature of the speech segment, D I,J represents a matrix of I rows and J columns, where I ≤ N and J ≤ M. It should be understood that the above examples should not be construed as limiting the present embodiment.

[0108] 103. Extract the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment.

[0109] In this embodiment, the step of "extracting the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment" may include:

[0110] Obtain the historical temporal features of the historical speech segment;

[0111] Extract the temporal feature information of the current speech segment based on the two-dimensional convolutional speech feature of the current speech segment and the historical temporal features of the historical speech segment;

[0112] Update the historical temporal features of the historical speech segment based on the two-dimensional convolutional speech feature of the current speech segment;

[0113] Take the next speech segment as the new current speech segment, and return to execute the step of obtaining the historical temporal features of the historical speech segment until the temporal feature information of each speech segment is obtained.

[0114] Among them, in the order of the speech segments, each speech segment is sequentially used as the currently processed speech segment. When the processed speech segment is the first speech segment, the temporal feature information of the current speech segment can be extracted only based on the two-dimensional convolutional speech feature of the current speech segment; or, the temporal feature information of the current speech segment can also be extracted based on the two-dimensional convolutional speech feature of the current speech segment and the historical temporal features of the historical speech segment, where the historical temporal features of the historical speech segment can be randomly generated or preset. After the temporal feature information of the first speech segment is extracted, the historical temporal features of the historical speech segment also need to be updated based on the two-dimensional convolutional speech feature of the first speech segment. When the processed speech segment is not the first speech segment, the temporal feature information of the current speech segment is extracted based on the two-dimensional convolutional speech feature of the current speech segment and the historical temporal features of the historical speech segment. After the temporal feature information of the current speech segment is extracted, the historical temporal features of the historical speech segment also need to be updated, and this process is continuously looped until the temporal feature information of each speech segment is obtained.

[0115] Optionally, in some embodiments, the step of "extracting the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment" may include:

[0116] Using a recurrent neural network, based on the two-dimensional convolutional speech features of each speech segment, the temporal feature information of each speech segment is extracted.

[0117] Among them, the recurrent neural network (RNN, Recurrent Neural Network) can be a long short-term memory network (LSTM, Long Short-Term Memory), or a two-layer gated recurrent unit network (GRU, Gatedrecurrent units), etc. The LSTM can selectively forget part of the historical data through three gate structures (input gate, forget gate, output gate), add part of the current input data, and finally integrate it into the current state and generate an output state. The LSTM is more suitable for extracting semantic features from temporal data and is often used to extract semantic features from context information in natural language processing tasks. The GRU is a type of recurrent neural network, and like the LSTM, it is also proposed to solve problems such as long-term memory and gradients in backpropagation. In the GRU model, there are only two gates, namely the update gate and the reset gate; the update gate is used to control the degree to which the state information of the previous moment is brought into the current state. The larger the value of the update gate, the more state information of the previous moment is brought in. The reset gate is used to control the degree of ignoring the state information of the previous moment. The smaller the value of the reset gate, the more it is ignored. In addition, the GRU has one less gate function than the LSTM, so the number of parameters is also less than that of the LSTM. Therefore, the overall training speed of the GRU is faster than that of the LSTM, which can greatly improve the training efficiency. Optionally, in some embodiments, see Figure 1c , by passing the spectral information (i.e., spectrogram) of each speech segment through the multi-layer convolutional layer and pooling layer of a two-dimensional convolutional neural network, the two-dimensional convolutional speech features of each speech segment can be obtained. Among them, the spectral information of the speech segment contains information in two dimensions: frequency and time. Using a recurrent neural network, based on the two-dimensional convolutional speech features of each speech segment, the temporal feature information of each speech segment can be extracted.

[0118] It can be understood that other machine learning methods can also be used in this embodiment to obtain the temporal feature information of speech segments, not limited to the method of artificial neural networks. For example, the temporal feature information of speech segments can also be described by a Markov model.

[0119] 104. Based on the temporal feature information of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category.

[0120] Among them, a classifier can be used to predict the attribute categories of each speech segment. Specifically, the classifier can be a support vector machine (SVM, Support Vector Machine), or a recurrent neural network, or a fully connected deep neural network (DNN, Deep Neual Networks), etc. This embodiment does not limit this.

[0121] Among them, the attributes of the user can be divided into at least two attribute categories. For example, the gender of the user can be divided into two attribute categories: male and female; and for the age of the user, it can be divided into four attribute categories: 0 - 9 years old, 10 - 14 years old, 14 - 18 years old, and over 18 years old according to actual needs.

[0122] Optionally, in some embodiments, the attribute categories include age range and gender; the step "Based on the temporal feature information of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category" may include:

[0123] Based on the temporal feature information of each speech segment, predict the probability that the corresponding age range of each speech segment is each preset age range, and the probability that the corresponding gender is each preset gender.

[0124] Among them, the preset genders can be divided into two categories: male and female; the preset age ranges can be divided according to actual needs, and this embodiment does not limit this. For example, in the specific scenario of online games, the preset age ranges can be divided into four preset age ranges: 0 - 9 years old, 10 - 14 years old, 14 - 18 years old, and over 18 years old.

[0125] For example, for the first speech segment The predicted probability of 0 - 9 years old is 0.2, the probability of 10 - 14 years old is 0.5, the probability of 14 - 18 years old is 0.2, and the probability of over 18 years old is 0.1. For the second speech segment The predicted probability of 0 - 9 years old is 0.1, the probability of 10 - 14 years old is 0.5, the probability of 14 - 18 years old is 0.3, and the probability of over 18 years old is 0.1...

[0126] 105. Based on the preset attribute categories, fuse the probabilities of the predicted speech segments to obtain the fused probabilities corresponding to each preset attribute category.

[0127] Optionally, in some embodiments, the step "Based on the preset attribute categories, fuse the probabilities of the predicted speech segments to obtain the fused probabilities corresponding to each preset attribute category" may include:

[0128] Perform a weighted average on the probabilities of the same preset attribute categories for each predicted speech segment to obtain the fused probability corresponding to each preset attribute category.

[0129] For example, there are J speech segments j represents the label of the speech segment in different time periods. The attributes of the user can be divided into I preset attribute categories, where i ∈ {1, 2, 3...I}, and i represents the label of different preset attribute categories. For each speech segment Finally, the result can be obtained indicating the probability of the speech segment for the j-th time period t j of the probability that the attribute category is each preset attribute category.

[0130] Among them, by smoothing the probability of each preset attribute category, the mean value of the probabilities of each speech segment in this preset attribute category can be taken to obtain the fused probability corresponding to each preset attribute category. That is, for the i-th preset attribute category, the smoothed probability is as shown in formula (3):

[0131]

[0132] Among them, represents the probability of the j-th speech segment in the i-th preset attribute category, is the probability of the i-th preset attribute category smoothed by J speech segments, that is, the fused probability corresponding to the i-th preset attribute category.

[0133] Optionally, in some embodiments, the attribute category includes age group and gender; the step of "fusing the probabilities of each predicted speech segment based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category" may include:

[0134] Fusing the probabilities of the corresponding age groups of each predicted speech segment based on the preset age group to obtain the fused probability corresponding to each preset age group;

[0135] Fusing the probabilities of the corresponding genders of each predicted speech segment based on the preset gender to obtain the fused probability corresponding to each preset gender.

[0136] For example, there are a total of 3 speech segments, and the preset age groups are divided into four preset age groups: 0 - 9 years old, 10 - 14 years old, 14 - 18 years old, and over 18 years old. For the first speech segment the predicted probability of 0 - 9 years old is 0.2, the probability of 10 - 14 years old is 0.5, the probability of 14 - 18 years old is 0.2, and the probability of over 18 years old is 0.1; for the second speech segment The predicted probabilities are 0.1 for the age range of 0 - 9 years old, 0.5 for 10 - 14 years old, 0.3 for 14 - 18 years old, and 0.1 for over 18 years old; for the third voice segment The predicted probabilities are 0.2 for the age range of 0 - 9 years old, 0.4 for 10 - 14 years old, 0.2 for 14 - 18 years old, and 0.2 for over 18 years old; then for the age range of 0 - 9 years old, the fused probability is (0.2 + 0.1 + 0.2) / 3, for 10 - 14 years old is (0.5 + 0.5 + 0.4) / 3, for 14 - 18 years old is (0.2 + 0.3 + 0.2) / 3, and for over 18 years old is (0.1 + 0.1 + 0.2) / 3.

[0137] For another example, there are a total of 2 voice segments, and the preset gender is divided into two categories: male and female. For the first voice segment The predicted probability that the user corresponding to the voice segment is female is 0.7, and male is 0.3. For the second voice segment The predicted probability that the user corresponding to the voice segment is female is 0.9, and male is 0.1. Then the fused probability for the target user being female is 0.8, and for male is 0.2.

[0138] 106. Based on the fused probabilities corresponding to each preset attribute category, determine the target attribute category corresponding to the target user from the preset attribute categories.

[0139] Optionally, in some embodiments, the step of "Based on the fused probabilities corresponding to each preset attribute category, determine the target attribute category corresponding to the target user from the preset attribute categories" may include:

[0140] Obtain the maximum fused probability among the fused probabilities corresponding to each preset attribute category;

[0141] Determine the preset attribute category corresponding to the maximum fused probability as the target attribute category corresponding to the target user.

[0142] Among them, based on the fused probabilities corresponding to each preset attribute category calculated in step 105 Select the maximum fused probability from them, and the preset attribute category corresponding to the maximum fused probability is the target attribute category, that is, calculate The i-th attribute category corresponding to the maximum fused probability is the target attribute category.

[0143] Optionally, in some embodiments, the attribute categories include age group and gender; the step of "determining the target attribute category corresponding to the target user from the preset attribute categories based on the fused probabilities corresponding to the respective preset attribute categories" may include:

[0144] Determining the target age group corresponding to the target user from the preset age groups based on the fused probabilities corresponding to the respective preset age groups;

[0145] Determining the target gender corresponding to the target user from the preset genders based on the fused probabilities corresponding to the respective preset genders.

[0146] Among them, the preset age group corresponding to the maximum fused probability can be selected as the target age group; the preset gender corresponding to the maximum fused probability can be selected as the target gender.

[0147] In some embodiments, as Figure 1d shown, using short-time orthogonal transformation, multiple speech segments of the target user are transformed from the time domain to the frequency domain to obtain the spectral information of each speech segment, that is, the spectrogram, and then the spectral information of each speech segment is subjected to feature extraction through a convolutional neural network and a recurrent neural network to obtain the temporal feature information of each speech segment; then, based on the temporal feature information of each speech segment, the probability that the attribute category of each speech segment is each preset attribute category can be predicted, and finally, the probabilities of the speech segments are smoothed. Specifically, based on the preset attribute categories, the probabilities of the predicted speech segments are fused to obtain the fused probabilities corresponding to the respective preset attribute categories, the maximum fused probability among the fused probabilities corresponding to the respective preset attribute categories is obtained, and the preset attribute category corresponding to the maximum fused probability is determined as the target attribute category corresponding to the target user.

[0148] It should be noted that the two-dimensional convolutional neural network and the recurrent neural network in this embodiment are trained by multiple pieces of training data, and the training data may include sample speech segments and the target attribute categories corresponding to the sample speech segments; the two-dimensional convolutional neural network and the recurrent neural network can be specifically trained by other devices and then provided to the user attribute recognition device, or, alternatively, can be trained by the user attribute recognition device itself, see Figure 1e .

[0149] If the user attribute recognition device trains itself, before the step of "performing convolutional feature extraction on each speech segment through a two-dimensional convolutional neural network to obtain the two-dimensional convolutional speech features of each speech segment", the user attribute recognition method may further include:

[0150] Obtaining training data, where the training data includes sample speech segments and the target attribute categories corresponding to the sample speech segments;

[0151] Performing convolutional feature extraction on the sample speech segment through a two-dimensional convolutional neural network to obtain the two-dimensional convolutional speech features of the sample speech segment;

[0152] Adopting a recurrent neural network, based on the two-dimensional convolutional speech features of the sample speech segment, to extract the temporal feature information of the sample speech segment;

[0153] Based on the temporal feature information of the sample speech segment, predicting the probability that the attribute category of the sample speech segment is each preset attribute category;

[0154] Adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probability that the predicted attribute category of the sample speech segment is the target attribute category meets a preset condition.

[0155] Optionally, in some embodiments, the step of "adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probability that the predicted attribute category of the sample speech segment is the target attribute category meets a preset condition" may include:

[0156] Calculating the loss value between the probability that the predicted attribute category of the sample speech segment is each preset attribute category and the true probability, wherein the true probability that the attribute category of the sample speech segment is the target attribute category is 1, and the true probability that the attribute category of the sample speech segment is other preset attribute categories except the target attribute category is 0;

[0157] Based on the loss value, adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the calculated loss value is less than a preset loss value.

[0158] Among them, the preset loss value can be set according to the actual situation, and this embodiment does not limit this. For this embodiment, the loss value between the probability that the predicted attribute category of the sample speech segment is each preset attribute category and the true probability can be calculated through a loss function, and this loss function can be cross entropy (CrossEntropy), etc. The formula of cross entropy is shown in formula (4):

[0159] L = ∑-p i log(p i ) (4)

[0160] Among them, L corresponds to the loss value between the probability that the predicted attribute category of the sample speech segment is each preset attribute category and the true probability, and p i is the probability of the sample speech segment in the i-th preset attribute category.

[0161] Optionally, if there is a requirement for the false alarm probability, a multiplicative constraint can be added to the above formula to reduce the false alarm rate, as shown in formula (5):

[0162] L mask = ∑ i -p i log(p i ) × ∑ t,q w tq δ(p q , max(p q ) p t ) (5)

[0163] Wherein, L mask is the loss value between the probability of the attribute category of the sample voice segment with multiplicative constraint being each preset attribute category and the true probability. δ(*) is the Dirac function. p q is the probability of the q-th attribute category. p q = {p1, p2,...}, p t is the one-hot encoded annotation value. w tq is the weight of classification, indicating the weight that the q-th attribute category being classified as the t-th attribute category by the classifier occupies in the loss function. Among them, the one-hot encoding uses an N-bit status register to encode N states. Each state has an independent register bit, and only one of these register bits is valid, that is, only one state can be valid. The one-hot encoding is the representation of categorical variables as binary vectors. It requires mapping the categorical values to integer values, and then each integer value is represented as a binary vector, which is all zero values except for the index of the integer, and it is marked as 1.

[0164] Optionally, in this embodiment, the sample voice segment is annotated to determine the target attribute category corresponding to the sample voice segment. The training process is to first predict the probability that the attribute category of the sample voice segment is each preset attribute category, and then use the backpropagation algorithm to adjust the parameters of the two-dimensional convolutional neural network and the recurrent neural network. Based on the loss value between the probability that the predicted attribute category of the sample voice segment is each preset attribute category and the true probability, the parameters of the two-dimensional convolutional neural network and the recurrent neural network are updated to make the probability that the predicted attribute category of the sample voice segment is the target attribute category approach the true probability 1, and the trained two-dimensional convolutional neural network and recurrent neural network are obtained. Specifically, the probability that the predicted attribute category of the sample voice segment is the target attribute category can be made higher than the preset probability, where the preset probability can be set according to the actual situation.

[0165] The user attribute recognition method provided in this embodiment has a shorter link compared with the solution of recognizing user attribute categories through images and videos, greatly reducing the amount of calculation. It can also effectively prevent the leakage of user privacy and improve the information security of the terminal. Moreover, with the wide application of intelligent devices, recording devices such as microphones are everywhere around everyone. Therefore, the deployment of the user attribute recognition device in this embodiment is relatively convenient. In addition, the recognition accuracy of this user attribute recognition device is relatively high, and the system robustness is relatively strong.

[0166] As can be seen from the above, this embodiment can obtain multiple speech segments of a target user with a temporal relationship, extract convolutional features from each speech segment to obtain two-dimensional convolutional speech features of each speech segment, extract temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category based on the temporal feature information of each speech segment, fuse the probabilities of each predicted speech segment based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category, and determine the target attribute category corresponding to the target user from the preset attribute categories based on the fused probability corresponding to each preset attribute category. By improving the method for recognizing the attribute category of the target user, this application can improve the operation speed and result accuracy of user attribute recognition.

[0167] Embodiment 2

[0168] According to the method described in the previous embodiment, the following will take the specific integration of this user attribute recognition device in the server as an example for further detailed description.

[0169] An embodiment of this application provides a user attribute recognition method. As Figure 2a shown, the specific process of this user attribute recognition method can be as follows:

[0170] 201. The server receives multiple speech segments of a target user sent by the terminal.

[0171] In this embodiment, the target user is the user for whom the attribute category needs to be obtained. Among them, the attribute category can be age, gender, etc., and this embodiment does not limit this.

[0172] 202. The server extracts convolutional features from each speech segment to obtain two-dimensional convolutional speech features of each speech segment.

[0173] In this embodiment, the step "The server extracts convolutional features from each speech segment to obtain two-dimensional convolutional speech features of each speech segment" may include:

[0174] Divide each speech segment into at least one speech frame;

[0175] Transform the speech frame from the time domain to the frequency domain to obtain the sub-spectrum information of the speech frame;

[0176] Fuse the sub-spectrum information of each speech frame in each speech segment to obtain the spectrum information of each speech segment;

[0177] Extract convolution features from the spectrum information of each speech segment to obtain the two-dimensional convolution speech features of each speech segment.

[0178] Among them, to transform the speech frame from the time domain to the frequency domain, the short-time discrete Fourier transform (STDFT, Short Time Discrete Fourier Transform) technology or the short-time discrete cosine transform technology (STDCT, Short Time Discrete Cosine Transform) can be used to orthogonally decompose the time-frequency of the speech frame signal to obtain the sub-spectrum information of the speech frame.

[0179] Among them, to fuse the sub-spectrum information of each speech frame in each speech segment, specifically, it can be to splice the sub-spectrum information of each speech frame in each speech segment to obtain the spectrum information of each speech segment. For example, splicing can be performed according to the timing relationship of the speech frames.

[0180] Optionally, in some embodiments, the step of "extracting convolution features from each speech segment to obtain the two-dimensional convolution speech features of each speech segment" may include:

[0181] Extract convolution features from each speech segment to obtain the two-dimensional convolution feature maps of each speech segment;

[0182] Perform a dimensionality reduction operation on the two-dimensional convolution feature maps of each speech segment to obtain the two-dimensional convolution speech features of each speech segment.

[0183] Among them, convolution features can be extracted from each speech segment through a neural network. The neural network can be a convolutional neural network, a visual geometry group network, a residual network, a densely connected convolutional network, etc. However, it should be understood that the neural network in this embodiment is not limited to the several types listed above. In addition, other machine learning methods can also be used to obtain the feature information of the speech segment, such as various acoustic features such as Mel spectrum and pitch, not limited to the method of artificial neural network.

[0184] Optionally, the dimensionality reduction of the two-dimensional convolution feature maps of each speech segment can be performed by pooling the two-dimensional convolution feature maps of each speech segment. To perform dimensionality reduction on the two-dimensional convolution feature maps of each speech segment.

[0185] 203. The server extracts the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment.

[0186] In this embodiment, the step of "extracting the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment" may include:

[0187] Obtain the historical temporal features of the historical speech segment;

[0188] Based on the two-dimensional convolutional speech features of the current speech segment and the historical temporal features of the historical speech segment, extract the temporal feature information of the current speech segment;

[0189] Based on the two-dimensional convolutional speech features of the current speech segment, update the historical temporal features of the historical speech segment;

[0190] Take the next speech segment as the new current speech segment, and return to execute the step of obtaining the historical temporal features of the historical speech segment until the temporal feature information of each speech segment is obtained.

[0191] Optionally, in some embodiments, the step of "extracting the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment" may include:

[0192] Use a recurrent neural network to extract the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment.

[0193] Among them, the recurrent neural network (RNN, Recurrent Neural Network) can be a long short-term memory network (LSTM, Long Short-Term Memory), or a gated recurrent unit network (GRU, Gatedrecurrent units), etc. It can be understood that other machine learning methods can also be used in this embodiment to obtain the temporal feature information of speech segments, not limited to the method of artificial neural network.

[0194] 204. The server predicts the probability that the attribute category of each speech segment is each preset attribute category based on the temporal feature information of each speech segment.

[0195] Among them, a classifier can be used to predict the attribute category of each speech segment. The classifier can specifically be a support vector machine, a fully connected deep neural network, etc. This embodiment does not limit this.

[0196] Among them, the attributes of the user can be divided into at least two attribute categories. For example, the gender of the user can be divided into two attribute categories: male and female; and for the age of the user, it can be divided into four attribute categories: 0 - 9 years old, 10 - 14 years old, 14 - 18 years old, and over 18 years old according to actual needs.

[0197] Optionally, in some embodiments, the attribute categories include age group and gender; the step of "predicting the probability that each voice segment belongs to each preset attribute category based on the temporal feature information of each voice segment" may include:

[0198] Predicting the probability that the age group corresponding to each voice segment belongs to each preset age group and the probability that the gender belongs to each preset gender based on the temporal feature information of each voice segment.

[0199] 205. The server fuses the probabilities of the predicted voice segments based on the preset attribute categories to obtain the fused probabilities corresponding to each preset attribute category.

[0200] Optionally, in some embodiments, the step of "fusing the probabilities of the predicted voice segments based on the preset attribute categories to obtain the fused probabilities corresponding to each preset attribute category" may include:

[0201] Performing a weighted average on the probabilities of the same preset attribute category of the predicted voice segments to obtain the fused probabilities corresponding to each preset attribute category.

[0202] Optionally, in some embodiments, the attribute categories include age group and gender; the step of "fusing the probabilities of the predicted voice segments based on the preset attribute categories to obtain the fused probabilities corresponding to each preset attribute category" may include:

[0203] Fusing the probabilities of the age groups corresponding to the predicted voice segments based on the preset age groups to obtain the fused probabilities corresponding to each preset age group;

[0204] Fusing the probabilities of the genders corresponding to the predicted voice segments based on the preset genders to obtain the fused probabilities corresponding to each preset gender.

[0205] 206. Based on the fused probabilities corresponding to each preset attribute category, determine the target attribute category of the target user from the preset attribute categories.

[0206] Optionally, in some embodiments, the step of "determining the target attribute category of the target user from the preset attribute categories based on the fused probabilities corresponding to each preset attribute category" may include:

[0207] Obtain the maximum fused probability among the fused probabilities corresponding to each preset attribute category;

[0208] Determine the preset attribute category corresponding to the maximum fused probability as the target attribute category of the target user.

[0209] Optionally, in some embodiments, the attribute category includes age group and gender; the step of "determining the target attribute category of the target user from the preset attribute categories based on the fused probabilities corresponding to each preset attribute category" may include:

[0210] Determine the target age group of the target user from the preset age groups based on the fused probabilities corresponding to each preset age group;

[0211] Determine the target gender of the target user from the preset genders based on the fused probabilities corresponding to each preset gender.

[0212] Among them, the preset age group corresponding to the maximum fused probability can be selected as the target age group; the preset gender corresponding to the maximum fused probability can be selected as the target gender.

[0213] 207. The server sends the target attribute category of the target user to the terminal.

[0214] Optionally, in this embodiment, refer to Figure 2b , specifically, audio information of the target user can be collected through the terminal. The audio information includes multiple voice segments of the target user. The voice activity detector can be used to denoise the audio information to obtain multiple voice segments of the target user. Perform discrete time-frequency domain transformation on each voice segment to transform each voice segment from the time domain to the frequency domain to obtain the spectral information of each voice segment. Extract features and reduce the dimension of the spectral information of each voice segment, and then input the obtained feature information of each voice segment into the classifier to predict the probability that the attribute category of each voice segment is each preset attribute category. Then, based on the preset attribute categories, fuse the probabilities of each predicted voice segment to obtain the fused probabilities corresponding to each preset attribute category. Among them, the attribute category of the target user can include age and gender. Based on the fused probabilities corresponding to each preset attribute category, determine the target attribute category of the target user from the preset attribute categories. The target attribute category can include the target gender and the target age.

[0215] For example, in the scenario of an online game, to protect minors, it is necessary to limit the daily login time of minors. For example, it can be restricted that underage users aged 12 and below can log in to the game for no more than 1 hour per day, and underage users aged 13 and above can log in to the game for no more than 2 hours per day. When logging in to the game, a game interface can pop up such asFigure 2c Describe the pop-up box shown. In online games, players sometimes communicate with their teammates through the in-game voice intercom function. The call recordings are saved bypassed and can be processed by VAD to obtain multiple voice segments of the player. These multiple voice segments are input into the user attribute recognition device of this embodiment; the spectral information of the voice segments, that is, the spectrogram, is extracted by the user attribute recognition device, and then the temporal feature information of the voice segments is obtained through a neural network. Based on the temporal feature information of each voice segment, the probability of the corresponding age group of each voice segment is predicted. Finally, by smoothing the probabilities of multiple voice segments, the age group with the highest probability value is selected as the age range of the player. If the player falls within the range of underage users aged 12 and below, a pop-up warning as shown in Figure 2d will be displayed on the game interface, and the usage time of the player will be restricted.

[0216] As can be seen from the above, this embodiment can receive multiple voice segments with temporal relationships of the target user sent by the terminal through the server. The server performs convolutional feature extraction on each voice segment to obtain the two-dimensional convolutional voice features of each voice segment. Based on the two-dimensional convolutional voice features of each voice segment, the temporal feature information of each voice segment is extracted. Based on the temporal feature information of each voice segment, the probability that the attribute category of each voice segment is each preset attribute category is predicted. Based on the preset attribute category, the probabilities of the predicted voice segments are fused to obtain the fused probabilities corresponding to each preset attribute category. Based on the fused probabilities corresponding to each preset attribute category, the target attribute category corresponding to the target user is determined from the preset attribute categories. The server sends the target attribute category corresponding to the target user to the terminal; this application can improve the operation speed and result accuracy of user attribute recognition by improving the method of identifying the attribute category of the target user.

[0217] Embodiment III

[0218] To better implement the above method, this embodiment of the application also provides a user attribute recognition device, as shown in Figure 3a . The user attribute recognition device may include an acquisition unit 301, a first extraction unit 302, a second extraction unit 303, a prediction unit 304, a fusion unit 305, and a determination unit 306, as follows:

[0219] (1) Acquisition unit 301;

[0220] The acquisition unit 301 is used to acquire multiple voice segments with temporal relationships of the target user.

[0221] (2) First extraction unit 302;

[0222] The first extraction unit 302 is configured to perform convolutional feature extraction on each voice segment to obtain two-dimensional convolutional voice features of each voice segment.

[0223] Optionally, in some embodiments of the present application, the first extraction unit 302 may include a division subunit 3021, a transformation subunit 3022, and a first extraction subunit 3023. Refer to Figure 3b as follows:

[0224] The division subunit 3021 is configured to divide each voice segment into at least one voice frame;

[0225] The transformation subunit 3022 is configured to transform the voice frame from the time domain to the frequency domain to obtain sub-spectrum information of the voice frame;

[0226] The first extraction subunit 3023 is configured to perform convolutional feature extraction on the sub-spectrum information of the voice frames of each voice segment to obtain two-dimensional convolutional voice features of each voice segment.

[0227] Optionally, in some embodiments of the present application, the first extraction subunit 3023 may specifically be configured to fuse the sub-spectrum information of the voice frames in each voice segment to obtain spectrum information of each voice segment; perform convolutional feature extraction on the spectrum information of each voice segment to obtain two-dimensional convolutional voice features of each voice segment.

[0228] Optionally, in some embodiments of the present application, the first extraction unit 302 may include a convolutional subunit 3024 and a dimensionality reduction subunit 3025. Refer to Figure 3c as follows:

[0229] The convolutional subunit 3024 is configured to perform convolutional feature extraction on each voice segment to obtain two-dimensional convolutional feature maps of each voice segment;

[0230] The dimensionality reduction subunit 3025 is configured to perform a dimensionality reduction operation on the two-dimensional convolutional feature maps of each voice segment to obtain two-dimensional convolutional voice features of each voice segment.

[0231] Optionally, in some embodiments of the present application, the first extraction unit 302 may specifically perform convolutional feature extraction on each voice segment through a two-dimensional convolutional neural network to obtain two-dimensional convolutional voice features of each voice segment.

[0232] (3) The second extraction unit 303;

[0233] The second extraction unit 303 is configured to extract temporal feature information of each voice segment based on the two-dimensional convolutional voice features of each voice segment.

[0234] Optionally, in some embodiments of the present application, the second extraction unit 303 may include an acquisition subunit 3031, a second extraction subunit 3032, an update subunit 3033, and a return subunit 3034, see Figure 3d , as follows:

[0235] The acquisition subunit 3031 is configured to acquire the historical timing features of the historical voice segment;

[0236] The second extraction subunit 3032 is configured to extract the timing feature information of the current voice segment based on the two-dimensional convolutional voice features of the current voice segment and the historical timing features of the historical voice segment;

[0237] The update subunit 3033 is configured to update the historical timing features of the historical voice segment based on the two-dimensional convolutional voice features of the current voice segment;

[0238] The return subunit 3034 is configured to use the next voice segment as the new current voice segment, and return to execute the step of acquiring the historical timing features of the historical voice segment until the timing feature information of each voice segment is obtained.

[0239] Optionally, in some embodiments of the present application, the second extraction unit 303 may specifically use a recurrent neural network to extract the timing feature information of each voice segment based on the two-dimensional convolutional voice features of each voice segment.

[0240] (4) Prediction unit 304;

[0241] The prediction unit 304 is configured to predict the probability that the attribute category of each voice segment is each preset attribute category based on the timing feature information of each voice segment.

[0242] Optionally, in some embodiments of the present application, the attribute categories include age group and gender; the prediction unit 304 may specifically be configured to predict the probability that the age group corresponding to each voice segment is each preset age group and the probability that the corresponding gender is each preset gender based on the timing feature information of each voice segment.

[0243] (5) Fusion unit 305;

[0244] The fusion unit 305 is configured to fuse the probabilities of each predicted voice segment based on the preset attribute categories to obtain the fused probabilities corresponding to each preset attribute category.

[0245] Optionally, in some embodiments of the present application, the attribute categories include age group and gender; the fusion unit 305 may specifically be configured to fuse the probabilities of the age groups corresponding to the predicted voice segments based on the preset age groups to obtain the fused probabilities corresponding to each preset age group; and fuse the probabilities of the genders corresponding to the predicted voice segments based on the preset gender to obtain the fused probabilities corresponding to each preset gender.

[0246] (6) Determination unit 306;

[0247] The determination unit 306 is configured to determine the target attribute category corresponding to the target user from the preset attribute categories based on the fused probabilities corresponding to each preset attribute category.

[0248] Optionally, in some embodiments of the present application, the attribute categories include age group and gender; the determination unit 306 may specifically be configured to determine the target age group corresponding to the target user from the preset age groups based on the fused probabilities corresponding to each preset age group; and determine the target gender corresponding to the target user from the preset genders based on the fused probabilities corresponding to each preset gender.

[0249] Optionally, in some embodiments of the present application, the user attribute recognition device may further include a training unit 307, see Figure 3e , as follows:

[0250] The training unit 307 is specifically configured to obtain training data, where the training data includes sample voice segments and the target attribute categories corresponding to the sample voice segments; extract convolutional features of the sample voice segments through a two-dimensional convolutional neural network to obtain two-dimensional convolutional voice features of the sample voice segments; use a recurrent neural network to extract temporal feature information of the sample voice segments based on the two-dimensional convolutional voice features of the sample voice segments; predict the probabilities that the attribute categories of the sample voice segments are each preset attribute category based on the temporal feature information of the sample voice segments; and adjust the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probabilities that the predicted attribute categories of the sample voice segments are the target attribute categories meet a preset condition.

[0251] Among them, optionally, in some embodiments, the step of "adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probabilities that the predicted attribute categories of the sample voice segments are the target attribute categories meet a preset condition" may specifically include:

[0252] Calculate the loss value between the probability of each preset attribute category of the predicted attribute category of the sample voice segment and the true probability, where the true probability that the attribute category of the sample voice segment is the target attribute category is 1, and the true probability that the attribute category of the sample voice segment is other preset attribute categories except the target attribute category is 0;

[0253] Based on the loss value, adjust the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the calculated loss value is less than the preset loss value.

[0254] As can be seen from the above, in this embodiment, the acquisition unit 301 can acquire multiple voice segments with a temporal relationship of the target user, the first extraction unit 302 can perform convolutional feature extraction on each voice segment to obtain the two-dimensional convolutional voice features of each voice segment, and based on the two-dimensional convolutional voice features of each voice segment, the second extraction unit 303 can extract the temporal feature information of each voice segment. Based on the temporal feature information of each voice segment, the prediction unit 304 can predict the probability that the attribute category of each voice segment is each preset attribute category. Based on the preset attribute category, the fusion unit 305 can fuse the probabilities of the predicted voice segments to obtain the fused probability corresponding to each preset attribute category. Based on the fused probability corresponding to each preset attribute category, the determination unit 306 can determine the target attribute category corresponding to the target user from the preset attribute categories; by improving the method for identifying the attribute category of the target user, this application can improve the operation speed and result accuracy of user attribute recognition.

[0255] Embodiment 4

[0256] The embodiment of the present application further provides an electronic device, as Figure 4 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present application. Specifically:

[0257] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 4 the structure of the electronic device shown in

[0258] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the electronic device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.

[0259] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the electronic device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0260] The electronic device further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0261] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0262] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0263] Obtain multiple speech segments with a temporal relationship of the target user, perform convolutional feature extraction on each speech segment to obtain the two-dimensional convolutional speech features of each speech segment, extract the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category based on the temporal feature information of each speech segment, fuse the probabilities of each predicted speech segment based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category, and determine the target attribute category corresponding to the target user from the preset attribute categories based on the fused probability corresponding to each preset attribute category.

[0264] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0265] As can be seen from the above, in this embodiment, multiple speech segments with a temporal relationship of the target user can be obtained, convolutional feature extraction is performed on each speech segment to obtain the two-dimensional convolutional speech features of each speech segment, the temporal feature information of each speech segment is extracted based on the two-dimensional convolutional speech features of each speech segment, the probability that the attribute category of each speech segment is each preset attribute category is predicted based on the temporal feature information of each speech segment, the probabilities of each predicted speech segment are fused based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category, and the target attribute category corresponding to the target user is determined from the preset attribute categories based on the fused probability corresponding to each preset attribute category; by improving the method for identifying the attribute category of the target user, the operation speed and result accuracy of user attribute recognition can be improved in this application.

[0266] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0267] Therefore, an embodiment of this application provides a storage medium in which multiple instructions are stored. These instructions can be loaded by a processor to execute the steps in any one of the user attribute recognition methods provided by the embodiments of this application. For example, the instructions can execute the following steps:

[0268] Obtain multiple speech segments with a temporal relationship of a target user, perform convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment, extract temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment, predict the probability that the attribute category of each speech segment is each preset attribute category based on the temporal feature information of each speech segment, fuse the probabilities of the predicted speech segments based on the preset attribute category to obtain the fused probability corresponding to each preset attribute category, and determine the target attribute category corresponding to the target user from the preset attribute categories based on the fused probability corresponding to each preset attribute category.

[0269] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.

[0270] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0271] Since the instructions stored in the storage medium can execute the steps in any of the user attribute recognition methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the user attribute recognition methods provided in the embodiments of the present application can be realized. For details, reference may be made to the previous embodiments and will not be elaborated herein.

[0272] The system involved in the embodiments of the present application may be a distributed system formed by connecting a client and multiple nodes (any form of electronic device accessing the network, such as a server, a terminal) in a form of network communication.

[0273] Taking the distributed system as a blockchain system as an example, refer to Figure 5 , Figure 5FIG. 0 is an optional structural diagram of the distributed system 100 provided by the embodiments of the present application when applied to a blockchain system, which is formed by multiple nodes 200 (any form of computing device connected to the network, such as a server or a user terminal) and a client 300. A peer-to-peer (P2P) network is formed among the nodes. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In a distributed system, any machine such as a server or a terminal can join and become a node. A node includes a hardware layer, an intermediate layer, an operating system layer, and an application layer. In this embodiment, information such as the trained two-dimensional convolutional neural network and recurrent neural network can be stored in the shared ledger of the blockchain system through the nodes. An electronic device (such as a terminal or a server) can obtain information such as the trained two-dimensional convolutional neural network and recurrent neural network based on the record data stored in the shared ledger.

[0274] See Figure 5 the functions of each node in the shown blockchain system, and the functions involved include:

[0275] 1) Routing, which is the basic function of a node and is used to support communication between nodes.

[0276] In addition to the routing function, a node can also have the following functions:

[0277] 2) Application, which is used to be deployed in the blockchain, implement specific services according to actual business requirements, form record data by recording data related to the implemented functions, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes verify the source and integrity of the record data successfully, they add the record data to the temporary block.

[0278] For example, the services implemented by the application include:

[0279] 2.1) Wallet, which is used to provide the function of conducting electronic currency transactions, including initiating a transaction (that is, sending the transaction record of the current transaction to other nodes in the blockchain system. After other nodes verify successfully, as a response to acknowledging the validity of the transaction, they deposit the record data of the transaction into the temporary block of the blockchain. Of course, the wallet also supports querying the remaining electronic currency in the electronic currency address;

[0280] 2.2) Shared ledger, which is used to provide functions such as storage, query, and modification of account data, send the record data of the operations on the account data to other nodes in the blockchain system. After other nodes verify its validity, as a response to acknowledging the validity of the account data, they deposit the record data into the temporary block, and can also send a confirmation to the node that initiated the operation.

[0281] 2.3) A smart contract is a computerized protocol that can execute the terms of a contract. It is implemented by code deployed on a shared ledger and executed when certain conditions are met. According to actual business requirements, the code is used to complete automated transactions, such as querying the logistics status of the goods purchased by a buyer and transferring the buyer's electronic currency to the merchant's address after the buyer signs for the goods. Of course, smart contracts are not limited to executing contracts for transactions, but can also execute contracts for processing received information.

[0282] 3) A blockchain includes a series of blocks that are sequentially connected in the order of generation. Once a new block is added to the blockchain, it will not be removed again. The block records the record data submitted by nodes in the blockchain system.

[0283] See Figure 6 , Figure 6 It is an optional schematic diagram of the block structure provided by the embodiments of this application. Each block includes the hash value of the transaction records stored in this block (the hash value of this block) and the hash value of the previous block. The blocks are connected through the hash values to form a blockchain. In addition, the block may also include information such as the timestamp when the block is generated. A blockchain is essentially a decentralized database, a string of data blocks associated using cryptographic methods. Each data block contains relevant information for verifying the validity of its information (anti-counterfeiting) and generating the next block.

[0284] The above has introduced in detail a user attribute recognition method and device provided by the embodiments of this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those skilled in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for identifying user attributes, characterized in that, Including: Obtaining multiple speech segments with a temporal relationship of a target user; Performing convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment; Based on the two-dimensional convolutional speech features of each speech segment, extracting temporal feature information of each speech segment; Based on the temporal feature information of each speech segment, predicting the probability that the attribute category of each speech segment is each preset attribute category; Based on the preset attribute categories, fusing the predicted probabilities of each speech segment to obtain the fused probabilities corresponding to each preset attribute category; Based on the fused probabilities corresponding to each preset attribute category, determining the target attribute category corresponding to the target user from the preset attribute categories; wherein, the attribute categories include age group and gender; Wherein, the extracting the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment includes: Obtaining the historical temporal features of historical speech segments; Based on the two-dimensional convolutional speech features of the current speech segment and the historical temporal features of the historical speech segment, extracting the temporal feature information of the current speech segment; Based on the two-dimensional convolutional speech features of the current speech segment, updating the historical temporal features of the historical speech segment; Taking the next speech segment as the new current speech segment, and returning to execute the step of obtaining the historical temporal features of the historical speech segment until the temporal feature information of each speech segment is obtained.

2. The method according to claim 1, characterized in that, The performing convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment includes: Dividing each speech segment into at least one speech frame; Transforming the speech frame from the time domain to the frequency domain to obtain the sub-spectrum information of the speech frame; Performing convolutional feature extraction on the sub-spectrum information of the speech frames of each speech segment to obtain two-dimensional convolutional speech features of each speech segment.

3. The method according to claim 2, wherein The performing convolutional feature extraction on the sub-spectrum information of the speech frames of each speech segment to obtain two-dimensional convolutional speech features of each speech segment includes: Fusing the sub-spectrum information of each speech frame in each speech segment to obtain the spectrum information of each speech segment; Performing convolutional feature extraction on the spectrum information of each speech segment to obtain two-dimensional convolutional speech features of each speech segment.

4. The method according to claim 1, wherein The performing convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment includes: Performing convolutional feature extraction on each speech segment to obtain two-dimensional convolutional feature maps of each speech segment; Performing a dimensionality reduction operation on the two-dimensional convolutional feature maps of each speech segment to obtain two-dimensional convolutional speech features of each speech segment.

5. The method according to claim 1, wherein The performing convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment includes: Performing convolutional feature extraction on each speech segment through a two-dimensional convolutional neural network to obtain two-dimensional convolutional speech features of each speech segment; The extracting the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment includes: Adopting a recurrent neural network to extract the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment.

6. The method according to claim 5, wherein Before extracting the convolutional features of each speech segment through a two-dimensional convolutional neural network to obtain the two-dimensional convolutional speech features of each speech segment, it further includes: Obtaining training data, where the training data includes sample speech segments and the target attribute categories corresponding to the sample speech segments; Extracting convolutional features of the sample speech segments through a two-dimensional convolutional neural network to obtain the two-dimensional convolutional speech features of the sample speech segments; Using a recurrent neural network to extract the temporal feature information of the sample speech segments based on the two-dimensional convolutional speech features of the sample speech segments; Predicting the probabilities that the attribute categories of the sample speech segments are each preset attribute category based on the temporal feature information of the sample speech segments; Adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probability that the predicted attribute category of the sample speech segment is the target attribute category meets a preset condition.

7. The method according to claim 6, characterized in that, The adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network so that the probability that the predicted attribute category of the sample speech segment is the target attribute category meets a preset condition includes: Calculating the loss value between the probabilities that the predicted attribute category of the sample speech segment is each preset attribute category and the true probabilities, where the true probability that the attribute category of the sample speech segment is the target attribute category is 1, and the true probability that the attribute category of the sample speech segment is other preset attribute categories except the target attribute category is 0; Adjusting the parameters of the two-dimensional convolutional neural network and the recurrent neural network based on the loss value so that the calculated loss value is less than a preset loss value.

8. The method according to claim 1, wherein The predicting the probabilities that the attribute categories of each speech segment are each preset attribute category based on the temporal feature information of each speech segment includes: Predicting the probabilities that the corresponding age groups of each speech segment are each preset age groups and the corresponding genders are each preset genders based on the temporal feature information of each speech segment; The fusing the probabilities of each predicted speech segment based on the preset attribute categories to obtain the fused probabilities corresponding to each preset attribute category includes: Fusing the probabilities of the corresponding age groups of each predicted speech segment based on the preset age groups to obtain the fused probabilities corresponding to each preset age group; Fusing the probabilities of the corresponding genders of each predicted speech segment based on the preset genders to obtain the fused probabilities corresponding to each preset gender; The determining the target attribute category of the target user from the preset attribute categories based on the fused probabilities corresponding to each preset attribute category includes: Determining the target age group of the target user from the preset age groups based on the fused probabilities corresponding to each preset age group; Determining the target gender of the target user from the preset genders based on the fused probabilities corresponding to each preset gender.

9. A user attribute recognition device, characterized in that, It includes: An obtaining unit for obtaining multiple speech segments with temporal relationships of a target user; The first extraction unit is configured to perform convolutional feature extraction on each speech segment to obtain two-dimensional convolutional speech features of each speech segment; The second extraction unit is configured to extract temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment. Wherein, the extracting the temporal feature information of each speech segment based on the two-dimensional convolutional speech features of each speech segment includes: obtaining historical temporal features of a historical speech segment; extracting temporal feature information of a current speech segment based on the two-dimensional convolutional speech features of the current speech segment and the historical temporal features of the historical speech segment; updating the historical temporal features of the historical speech segment based on the two-dimensional convolutional speech features of the current speech segment; taking the next speech segment as the new current speech segment, and returning to execute the step of obtaining the historical temporal features of the historical speech segment until the temporal feature information of each speech segment is obtained; The prediction unit is configured to predict, based on the temporal feature information of each speech segment, the probability that the attribute category of each speech segment is each preset attribute category; The fusion unit is configured to fuse the probabilities of the predicted speech segments based on the preset attribute categories to obtain the fused probabilities corresponding to the preset attribute categories; The determination unit is configured to determine, based on the fused probabilities corresponding to the preset attribute categories, the target attribute category corresponding to the target user from the preset attribute categories; wherein, the attribute categories include age group and gender.

10. An electronic device, characterized in that, It includes a memory and a processor; the memory stores multiple instructions, and the processor loads the instructions to execute the steps in the user attribute recognition method according to any one of claims 1 to 8.

11. A storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the steps in the user attribute recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for realizing gender recognition based on intelligent voice dialogue

    CN109378007A

  • Physical sign data identification method and device, electronic equipment and storage medium

    CN110619889A