Method, device and electronic equipment for constructing user portrait
By using multimodal data fusion processing to extract voiceprint, body posture, and facial expression features, the problem of low accuracy in user profiling in existing technologies is solved, and higher-precision user profiling is achieved.
Patent Information
- Application Number
- CN202211739136.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In existing technologies, user profile construction mainly relies on single text data, resulting in low accuracy.
By acquiring multimodal data of target users, including audio data, video data, and image data, voiceprint features, body features, and facial expression features are extracted and fused with text data. User profiles are then constructed using clustering algorithms, genetic algorithms, or neural network algorithms.
It improved the accuracy of user profiles and enhanced the precision of user profile construction.
Smart Images

Figure CN116150415B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus and electronic device for constructing user profiles. Background Technology
[0002] Big data technology is an information processing technology that takes all the data resources of any system as its object and discovers the correlations between the data. It is currently widely used in many fields such as targeted advertising and personalized user services.
[0003] User profiling, as an important application of big data technology, aims to establish user attribute tags in multiple dimensions to outline user characteristics. This allows for subsequent analysis of user preferences based on these characteristics, thereby providing users with more efficient and targeted information pushes or user experiences that are closer to their personal habits.
[0004] Therefore, how to accurately construct user profiles is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for constructing user profiles, which can accurately construct user profiles, thereby improving the accuracy of the constructed user profiles.
[0006] This application provides a method for constructing user profiles, including:
[0007] Obtain multimodal data corresponding to the target user, wherein the multimodal data includes at least two of the following: audio data, video data, image data, or text data.
[0008] Determine the target text data corresponding to the multimodal data. The target text data includes the transformed text data obtained after transforming the non-text data in the multimodal data, or the target text data includes the transformed text data and the text data.
[0009] Based on the target text data and the multimodal data, a user profile corresponding to the target user is constructed.
[0010] According to a user profile construction method provided in this application, when the multimodal data includes the audio data, the step of constructing a user profile corresponding to the target user based on the target text data and the multimodal data includes:
[0011] Extract the voiceprint features of the target user from the audio data.
[0012] The target text data and the voiceprint features are fused together, and a user profile corresponding to the target user is constructed based on the fusion result.
[0013] According to the user profile construction method provided in this application, the method further includes:
[0014] The audio data is processed by a transcription algorithm to perform modal type conversion, resulting in first converted text data, which is the text data corresponding to the audio data.
[0015] According to a user profile construction method provided in this application, when the multimodal data includes the video data, the step of constructing a user profile corresponding to the target user based on the target text data and the multimodal data includes:
[0016] The body features of the target user are extracted from the video data, including action features and posture features.
[0017] The target text data and the body features are fused together, and a user profile corresponding to the target user is constructed based on the fusion result.
[0018] According to the user profile construction method provided in this application, the method further includes:
[0019] The video data is subjected to modality conversion processing to obtain the first audio data and the first image data corresponding to the video data.
[0020] A transcription algorithm is used to perform modal type conversion on the first audio data to obtain second converted text data corresponding to the first audio data; and an image analysis algorithm is used to perform modal type conversion on the first image data to obtain third converted text data corresponding to the first image data. The second converted text data and the third converted text data constitute the converted text data corresponding to the video data.
[0021] According to a user profile construction method provided in this application, when the multimodal data includes the image data, the step of constructing a user profile corresponding to the target user based on the target text data and the multimodal data includes:
[0022] Extract the facial expression features of the target user from the image data.
[0023] The target text data and the facial expression features are fused together, and a user profile corresponding to the target user is constructed based on the fusion result.
[0024] According to the user profile construction method provided in this application, the step of constructing a user profile corresponding to a target user based on the fusion result includes:
[0025] The fusion results are processed using clustering algorithms, genetic algorithms, or neural network algorithms to construct a user profile corresponding to the target user.
[0026] This application also provides a user profile construction apparatus, comprising:
[0027] The acquisition unit is used to acquire multimodal data corresponding to the target user, wherein the multimodal data includes at least two of the following: audio data, video data, image data, or text data.
[0028] The processing unit is configured to determine the target text data corresponding to the multimodal data, wherein the target text data includes converted text data obtained after converting non-text data in the multimodal data, or the target text data includes the converted text data and the text data.
[0029] The construction unit is used to construct a user profile corresponding to the target user based on the target text data and the multimodal data.
[0030] According to the user profile construction apparatus provided in this application, when the multimodal data includes the audio data, the construction unit is specifically used for:
[0031] Extract the voiceprint features of the target user from the audio data; fuse the target text data and the voiceprint features, and construct a user profile corresponding to the target user based on the fusion result.
[0032] According to the user profile construction apparatus provided in this application, the processing unit is further configured to perform modal type conversion processing on the audio data using a transcription algorithm to obtain first converted text data, wherein the first converted text data is the converted text data corresponding to the audio data.
[0033] According to the user profile construction apparatus provided in this application, when the multimodal data includes the video data, the construction unit is specifically used for:
[0034] The body features of the target user are extracted from the video data, including action features and posture features; the target text data and the body features are fused, and a user profile corresponding to the target user is constructed based on the fusion result.
[0035] According to the user profile construction apparatus provided in this application, the processing unit is further configured to:
[0036] The video data is subjected to modality conversion processing to obtain first audio data and first image data corresponding to the video data; the first audio data is subjected to modality conversion processing using a transcription algorithm to obtain second converted text data corresponding to the first audio data; and the first image data is subjected to modality conversion processing using an image analysis algorithm to obtain third converted text data corresponding to the first image data. The second converted text data and the third converted text data constitute the converted text data corresponding to the video data.
[0037] According to the user profile construction apparatus provided in this application, when the multimodal data includes the image data, the construction unit is specifically used for:
[0038] The facial expression features of the target user are extracted from the image data; the target text data and the facial expression features are fused, and a user profile corresponding to the target user is constructed based on the fusion result.
[0039] According to the user profile construction apparatus provided in this application, the construction unit is specifically used to: process the fusion result using a clustering algorithm, a genetic algorithm, or a neural network algorithm to construct a user profile corresponding to the target user.
[0040] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the user profile construction method as described above.
[0041] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the user profile construction method as described above.
[0042] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the user profile construction method as described above.
[0043] The user profile construction method, apparatus, and electronic device provided in this application acquire multimodal data corresponding to a target user, including at least two of audio data, video data, image data, or text data; and based on the multimodal data corresponding to the target user, determine the target text data corresponding to the multimodal data, the target text data including converted text data obtained after converting non-text data in the multimodal data, or the target text data including converted text data and text data; and then construct a user profile corresponding to the target user based on the target text data and multimodal data. Compared with constructing a user profile using a single text data, this method can accurately construct the user profile, thereby improving the accuracy of the constructed user profile. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 A flowchart illustrating a method for constructing a user profile as provided in an embodiment of this application;
[0046] Figure 2 A schematic diagram illustrating the construction of a user profile as provided in an embodiment of this application;
[0047] Figure 3 A flowchart illustrating another method for constructing a user profile provided in an embodiment of this application;
[0048] Figure 4 This is a schematic diagram of the structure of the user profile construction apparatus provided in the embodiments of this application;
[0049] Figure 5 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0051] In the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. In the textual description of this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0052] The technical solutions provided in this application can be applied to many scenarios such as advertising push, personalized user services and improvements. Taking the personalized recommendation scenario in the personalized user service and improvement scenario as an example, if a user profile can be accurately constructed, the user profile tags can be accurately predicted based on the constructed user profile, and then personalized recommendations can be made in a targeted manner according to the user profile tags. Therefore, how to accurately construct a user profile is crucial.
[0053] In existing technologies, user profiling primarily involves acquiring users' text data from social media platforms and constructing user profiles based on this text data. The resulting user profiles are then used to predict user tags. However, given the limited data types used for profiling and the neglect of other data types, the existing methods often result in low accuracy of the constructed user profiles.
[0054] In order to accurately construct user profiles and improve their accuracy, considering that in addition to text data, audio data, video data, or image data can also be used to characterize user attributes, it is advisable to construct user profiles based on at least two of the following multimodal data: audio data, video data, image data, or text data.
[0055] Therefore, based on the above technical concept, this application provides a method for constructing a user profile, which involves acquiring multimodal data corresponding to a target user, wherein the multimodal data includes at least two of audio data, video data, image data, or text data; and determining target text data corresponding to the multimodal data based on the multimodal data corresponding to the target user, wherein the target text data includes converted text data obtained after converting non-text data in the multimodal data, or the target text data includes converted text data and text data; and then constructing a user profile corresponding to the target user based on the target text data and the multimodal data.
[0056] Based on the above description, when constructing a user profile, multimodal data is acquired, including at least two of the following: audio data, video data, image data, or text data. The user profile corresponding to the target user is constructed based on the target text data and the multimodal data. Compared with constructing a user profile using a single text data, this method can accurately construct the user profile, thereby improving the accuracy of the constructed user profile.
[0057] To facilitate understanding of the user profile construction method provided in this application, the following detailed description will be provided through several specific embodiments. It is understood that these specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0058] Example 1
[0059] Figure 1 This is a flowchart illustrating a method for constructing a user profile according to an embodiment of this application. This method can be executed by software and / or hardware devices. For an example, please refer to [link to example]. Figure 1 As shown, the method for constructing this user profile may include:
[0060] S101. Obtain multimodal data corresponding to the target user. Multimodal data includes at least two of the following: audio data, video data, image data, or text data.
[0061] For example, multimodal data includes at least two of the following: audio data, video data, image data, or text data. In other words, multimodal data can include any two types of data, such as audio data and video data; it can also include any three types of data, such as audio data, video data, image data, or text data; or it can include audio data, video data, image data, and text data, depending on actual needs.
[0062] For example, when acquiring multimodal data corresponding to a target user, web crawlers can be used on different social media platforms to obtain the required multimodal data. A web crawler is a program that uses a computer to simulate a user's web browsing behavior, sending requests to websites to retrieve resources and obtaining the data returned by the websites.
[0063] Taking Weibo as an example, when performing web scraping on Weibo, one can open a user's Weibo homepage and use a computer program to scrape the user's posts published in November to extract multimodal data corresponding to the user. For instance, after scraping the multimodal data, it can be stored in a database for subsequent user profile construction. Furthermore, the data type of the scraped multimodal data can be labeled, allowing for the combination of these data types to construct user profiles. Of course, the data source of the multimodal data can also be labeled, indicating which social media platform the multimodal data was obtained from. For example, multimodal data obtained from Weibo or Zhihu.
[0064] After obtaining the multimodal data corresponding to the target user, the following step S102 can be executed:
[0065] S102. Determine the target text data corresponding to the multimodal data. The target text data includes the transformed text data obtained after transforming the non-text data in the multimodal data, or the target text data includes the transformed text data and the text data.
[0066] For example, when determining the target text data corresponding to multimodal data, the following two possible scenarios can be included:
[0067] In one possible scenario, the multimodal data does not include text data, that is, it only includes at least two types of non-text data such as audio data, video data, or image data. In this case, when determining the target text data corresponding to the multimodal data, it is necessary to perform conversion processing on each type of data in the multimodal data to obtain the corresponding converted text data.
[0068] In this possible scenario, the target text data corresponding to the multimodal data only includes the transformed text data obtained after transforming the non-text data in the multimodal data.
[0069] In another possible scenario, multimodal data includes not only text data but also at least one type of non-text data, such as audio data, video data, or image data. In this case, when determining the target text data corresponding to the multimodal data, it is not necessary to convert the text data. Instead, it is only necessary to convert each type of non-text data to obtain the corresponding converted text data.
[0070] In this possible scenario, the target text data corresponding to multimodal data includes not only the transformed text data obtained after converting non-text data, but also the text data in the multimodal data.
[0071] For example, in any of the above possible scenarios, when the multimodal data includes audio data, a transcription algorithm can be used to perform modality type conversion on the audio data to obtain the first converted text data corresponding to the audio data. This first converted text data is the converted text data obtained after converting the audio data.
[0072] When multimodal data includes video data, the moviepy package in Python can be used to perform modality conversion on the video data to obtain the first audio data and the first image data corresponding to the video data. A transcription algorithm is then used to perform modality conversion on the first audio data to obtain the second converted text data corresponding to the first audio data. Finally, an image analysis algorithm is used to perform modality conversion on the first image data to obtain the third converted text data corresponding to the first image data. The second and third converted text data are the converted text data obtained after converting the video data.
[0073] When multimodal data includes image data, image analysis algorithms can be used to perform modality conversion on the image data to obtain the fourth converted text data corresponding to the image data. This fourth converted text data is the converted text data obtained after converting the image data.
[0074] Based on the above description, after determining the target text data corresponding to the multimodal data, the following step S103 can be executed:
[0075] S103. Based on the target text data and multimodal data, construct a user profile corresponding to the target user.
[0076] For example, based on target text data and multimodal data, a user profile corresponding to the target user can be constructed, which can be described in combination with a variety of possible scenarios:
[0077] In one possible scenario, when multimodal data includes audio data, considering that audio data can be well used to extract the voiceprint features of the target user, when constructing a user profile corresponding to the target user based on the target text data and multimodal data, the voiceprint features of the target user can be extracted from the audio data first; and the target text data and voiceprint features can be fused, with the fusion result including the voiceprint features of the target user. In this way, by fusing voiceprint features into the target text data, the user profile corresponding to the target user can be accurately constructed based on the fusion result, thereby improving the accuracy of the constructed user profile.
[0078] For example, when extracting the voiceprint features of a target user from audio data, some existing voiceprint feature extraction algorithms or deep network-based voiceprint feature extraction models can be used to extract the voiceprint features of the target user from the audio data. Here, the embodiments of this application will not be described in detail.
[0079] In another possible scenario, where multimodal data includes video data, considering that video data can be well used to extract the target user's body features, including action and posture features, when constructing a user profile based on the target text data and multimodal data, the target user's body features can be extracted from the video data first. Then, the target text data and body features are fused, and the fusion result includes the target user's body features. By fusing body features into the target text data, the user profile can be accurately constructed based on the fusion result, thereby improving the accuracy of the constructed user profile.
[0080] For example, when extracting the body features of a target user from video data, some existing body feature extraction algorithms or body feature extraction models based on deep networks can be used to extract the body features of the target user from the video data. Here, the embodiments of this application will not be described in detail.
[0081] In another possible scenario, where multimodal data includes image data, considering that image data can be well used to extract facial expression features of the target user, when constructing a user profile corresponding to the target user based on the target text data and multimodal data, the facial expression features of the target user can be extracted from the image data first. Then, the target text data and facial expression features are fused, and the fusion result includes the facial expression features of the target user. In this way, by fusing facial expression features into the target text data, the user profile corresponding to the target user can be accurately constructed based on the fusion result, thereby improving the accuracy of the constructed user profile.
[0082] For example, when extracting facial expression features of a target user from image data, some existing facial expression feature extraction algorithms or deep network-based facial expression feature extraction models can be used to extract facial expression features of the target user from image data. Here, the embodiments of this application will not be described in detail.
[0083] It should be noted that the above description of constructing user profiles for target users based on target text data and multimodal data only refers to constructing user profiles by combining audio data, video data, or image data with the target text data. Alternatively, user profiles can be constructed by combining any two or three of the following: audio data, video data, and image data. For example, by combining audio and video data, voiceprint features of the target user can be extracted from the audio data, and body posture features can be extracted from the video data. The target text data, voice features, and body posture features can then be fused, and a user profile corresponding to the target user can be constructed based on the fusion result. Another example is the combination of audio data, video data, and image data. See also: Figure 2 As shown, Figure 2 This is a schematic diagram of a user profile construction method provided in an embodiment of this application. It can extract the voiceprint features of the target user from audio data, the body posture features of the target user from video data, and the facial expression features of the target user from image data. The target text data, voice features, body posture features, and facial expression features are fused together, and a user profile corresponding to the target user is constructed based on the fusion result. The specific settings can be configured according to actual needs.
[0084] For example, when constructing a user profile corresponding to a target user based on the fusion result, clustering algorithms, genetic algorithms, or neural network algorithms can be used to process the fusion result to construct the user profile corresponding to the target user.
[0085] As can be seen from the embodiments of this application, when constructing a user profile, multimodal data corresponding to the target user is obtained. The multimodal data includes at least two of the following: audio data, video data, image data, or text data. Based on the multimodal data corresponding to the target user, target text data corresponding to the multimodal data is determined. The target text data includes converted text data obtained after converting non-text data in the multimodal data, or the target text data includes converted text data and text data. Then, the user profile corresponding to the target user is constructed based on the target text data and the multimodal data. Compared with constructing a user profile using a single text data, the user profile can be constructed more accurately, thereby improving the accuracy of the constructed user profile.
[0086] To facilitate understanding of the user profile construction method provided in this application embodiment, a specific embodiment will be used to describe the user profile construction method provided in this application embodiment below.
[0087] Example 2
[0088] When building a user profile for a specific user, you can first use web crawlers to extract multimodal data from social media platforms such as Weibo or Zhihu. This multimodal data can then be categorized by its modality. For an example, see [link to example]. Figure 3 As shown, Figure 3 The flowchart of another user profile construction method provided in this application embodiment can be shown first. It can be determined whether the multimodal data includes video data. If the multimodal data includes video data, the video data is processed by Moviepy to obtain its corresponding image data and audio data. The image data is then processed by an image analysis algorithm to obtain converted text data after image data conversion. At the same time, the audio data is processed by a transcription algorithm to obtain converted text data after audio data conversion. The converted text data after image data conversion and the converted text data after audio data conversion are the converted text data obtained by video data modality type conversion processing.
[0089] If the multimodal data does not include video data, then it is determined whether the multimodal data includes image data. If it does, an image analysis algorithm is used to process the image data to obtain converted text data. Conversely, if the multimodal data does not include image data, it is further determined whether the multimodal data includes audio data. If it does, a transcription algorithm is used to process the audio data to obtain converted text data. Conversely, if the multimodal data does not include audio data, it is determined that the multimodal data includes text data. Assuming the multimodal data includes audio data, video data, image data, and text data, the target text data corresponding to the user includes converted text data from audio data, converted text data from video data, converted text data from image data, and text data.
[0090] After obtaining the target text data corresponding to the user, the target text data can be fused with the body features extracted from video data, the voiceprint features extracted from audio data, and the facial expression features extracted from image data. Then, clustering algorithms, genetic algorithms, or neural network algorithms are used to process the fusion results to construct the user profile corresponding to the user. This allows for accurate user profile construction and improves the accuracy of the constructed user profile.
[0091] The user profile construction apparatus provided in this application will be described below. The user profile construction apparatus described below can be referred to in correspondence with the user profile construction method described above.
[0092] Figure 4This is a schematic diagram of the user profile building apparatus provided in the embodiments of this application. For example, please refer to [link / reference]. Figure 4 As shown, the user profile building apparatus 40 may include:
[0093] The acquisition unit 401 is used to acquire multimodal data corresponding to the target user. The multimodal data includes at least two of the following: audio data, video data, image data, or text data.
[0094] Processing unit 402 is used to determine the target text data corresponding to the multimodal data. The target text data includes the transformed text data obtained after transforming the non-text data in the multimodal data, or the target text data includes the transformed text data and the text data.
[0095] Construction unit 403 is used to construct user profiles corresponding to target users based on target text data and multimodal data.
[0096] Optionally, when the multimodal data includes audio data, construction unit 403 is specifically used for:
[0097] Extract the voiceprint features of the target user from the audio data; fuse the target text data and voiceprint features, and construct a user profile corresponding to the target user based on the fusion result.
[0098] Optionally, the processing unit 402 is further configured to perform modal type conversion processing on the audio data using a transcription algorithm to obtain first converted text data, wherein the first converted text data is the transposed text data corresponding to the audio data.
[0099] Optionally, when the multimodal data includes video data, construction unit 403 is specifically used for:
[0100] The body features of the target user are extracted from the video data, including action features and posture features; the target text data and body features are fused, and a user profile corresponding to the target user is constructed based on the fusion result.
[0101] Optionally, the processing unit 402 is also used for:
[0102] The video data is processed by modality conversion to obtain the first audio data and the first image data corresponding to the video data; the first audio data is processed by modality conversion using a transcription algorithm to obtain the second converted text data corresponding to the first audio data; and the first image data is processed by modality conversion using an image analysis algorithm to obtain the third converted text data corresponding to the first image data. The second converted text data and the third converted text data constitute the converted text data corresponding to the video data.
[0103] Optionally, when the multimodal data includes image data, the construction unit 403 is specifically used for:
[0104] Extract facial expression features of the target user from image data; fuse the target text data and facial expression features, and construct a user profile corresponding to the target user based on the fusion result.
[0105] Optionally, the construction unit 403 is specifically used to: process the fusion result using a clustering algorithm, a genetic algorithm, or a neural network algorithm to construct a user profile corresponding to the target user.
[0106] The user profile construction apparatus 40 provided in this application embodiment can execute the technical solution of the user profile construction method in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the user profile construction method. Please refer to the implementation principle and beneficial effects of the user profile construction method. It will not be repeated here.
[0107] Figure 5 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of this application, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 550, wherein the processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 550. The processor 510 can call logical instructions in the memory 530 to execute a user profile construction method. This method includes: acquiring multimodal data corresponding to a target user, where the multimodal data includes at least two of audio data, video data, image data, or text data; determining target text data corresponding to the multimodal data, where the target text data includes converted text data obtained after converting non-text data from the multimodal data, or the target text data includes both converted text data and text data; and constructing a user profile corresponding to the target user based on the target text data and the multimodal data.
[0108] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0109] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the user profile construction method provided by the above methods. The method includes: acquiring multimodal data corresponding to a target user, wherein the multimodal data includes at least two of audio data, video data, image data, or text data; determining target text data corresponding to the multimodal data, wherein the target text data includes converted text data obtained after converting non-text data in the multimodal data, or the target text data includes converted text data and text data; and constructing a user profile corresponding to the target user based on the target text data and the multimodal data.
[0110] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a method for constructing a user profile provided by the methods described above. The method includes: acquiring multimodal data corresponding to a target user, wherein the multimodal data includes at least two of audio data, video data, image data, or text data; determining target text data corresponding to the multimodal data, wherein the target text data includes converted text data obtained after converting non-text data in the multimodal data, or the target text data includes converted text data and text data; and constructing a user profile corresponding to the target user based on the target text data and the multimodal data.
[0111] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for constructing a user profile, characterized in that, The method comprises: obtaining multi-modal data corresponding to a target user, the multi-modal data comprising at least two of audio data, video data, image data, or text data; determining target text data corresponding to the multi-modal data, the target text data comprising converted text data obtained by converting non-text data in the multi-modal data, or the target text data comprising the converted text data and the text data; based on the target text data and the multi-modal data, constructing a user portrait corresponding to the target user; wherein, in the case that the multi-modal data comprises the audio data, the constructing of the user portrait corresponding to the target user based on the target text data and the multi-modal data comprises: extracting a voiceprint feature of the target user from the audio data; fusing the target text data and the voiceprint feature, and constructing a user portrait corresponding to the target user based on the fusion result; in the case that the multi-modal data comprises the video data, the constructing of the user portrait corresponding to the target user based on the target text data and the multi-modal data comprises: extracting a body language feature of the target user from the video data, the body language feature comprising a motion feature and a posture feature; fusing the target text data and the body language feature, and constructing a user portrait corresponding to the target user based on the fusion result.
2. The user profiling method of claim 1, wherein, The method further comprises: performing modal type conversion processing on the audio data using a transcription algorithm to obtain first converted text data, the first converted text data being text data corresponding to the audio data.
3. The user profiling method of claim 1, wherein, The method further comprises: performing modal type conversion processing on the video data to obtain first audio data and first image data corresponding to the video data; performing modal type conversion processing on the first audio data using a transcription algorithm to obtain second converted text data corresponding to the first audio data, and performing modal type conversion processing on the first image data using an image analysis algorithm to obtain third converted text data corresponding to the first image data, the second converted text data and the third converted text data constituting converted text data corresponding to the video data.
4. The method of claim 1-3, wherein, in the case that the multi-modal data comprises the image data, the constructing of the user portrait corresponding to the target user based on the target text data and the multi-modal data comprises: extracting a facial expression feature of the target user from the image data; fusing the target text data and the facial expression feature, and constructing a user portrait corresponding to the target user based on the fusion result.
5. The method of constructing a user profile according to claim 1 or 2, wherein, The constructing of the user portrait corresponding to the target user based on the fusion result comprises: processing the fusion result using a clustering algorithm, a genetic algorithm, or a neural network algorithm to construct the user portrait corresponding to the target user.
6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the user portrait construction method of any one of claims 1 to 5.
7. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the user portrait construction method of any one of claims 1 to 5.
8. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method for constructing a user portrait according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method, device and equipment for generating customer portrait
CN113344067A
Information recommendation method and device based on multi-modal feature fusion and processor
CN114218488A