Financial user portrait prediction method and device, electronic equipment and storage medium
By using multimodal analysis methods combining video and text in financial user profiling prediction, the problems of low accuracy and poor timeliness in traditional methods are solved, enabling timely and accurate assessment of user assets during video communication and ensuring user privacy and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN BANK CO LTD
- Filing Date
- 2023-08-15
- Publication Date
- 2026-07-31
AI Technical Summary
In the financial sector, traditional financial user profiling prediction methods suffer from low accuracy and poor timeliness, making it difficult to obtain accurate user asset information in a timely manner during video communication, and resulting in a single source of information.
By acquiring the background area, text description, and image set of the user's video, a pre-set network model is used for recognition and analysis, including a video encoder, a text encoder, a multimodal joint encoder, and a text decoder, to output a text description of the financial user, avoiding the use of facial information to ensure privacy.
It enables accurate user profiling based on real-time information about the user's surrounding environment in online video scenarios, improving the timeliness and accuracy of user asset assessment and reducing the risk of personal information leakage.
Smart Images

Figure CN117079022B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device and storage medium for predicting financial user profiles. Background Technology
[0002] In the financial sector, financial institutions recommend suitable financial products to customers through communication. In traditional financial services, customer asset information is highly sensitive and, in principle, not visible to sales personnel at any stage of a financial transaction, such as when selling wealth management products. On the other hand, sales personnel, however, desire more information about customer assets to provide appropriate sales services.
[0003] Typically, manual rules are set up to predict a customer's asset potential level, such as "low / medium / high / very high potential level," and then customers who meet the level requirements are selected as candidates. This approach is relatively easy to implement, but its accuracy is low, and recommendations usually need to be made in real time during video communication, making the predictions from manual rules less timely. In addition, while video communication can solve the communication efficiency problem to some extent, it still suffers from the problem of a single information source, making it difficult for marketers to obtain accurate user profiles in a short period of time, resulting in a low accuracy rate for product recommendations. Summary of the Invention
[0004] This application provides a financial user profile prediction method, device, electronic device, and storage medium. It can identify and analyze the user's surroundings and background environment in real time in the context of online video, and obtain a corresponding description of the financial user profile while ensuring personal security information.
[0005] In a first aspect, embodiments of this application provide a method for predicting financial user profiles, the method comprising:
[0006] Acquire training video data, which includes: user video samples of background regions extracted from complete user videos, text descriptions corresponding to the user video samples, text classifications of the user video samples, and image sets corresponding to the user video samples, with the user video samples being classified based on the image sets.
[0007] The preset network model is trained using the video data to be trained to obtain the target network model;
[0008] The target network model is used to process the video to be processed to obtain the text description corresponding to the video to be processed.
[0009] The financial user profile corresponding to the video to be processed is determined by analyzing the text description corresponding to the video to be processed.
[0010] In one optional implementation, the step of processing the video to be processed based on the target network model to obtain a text description corresponding to the video to be processed includes:
[0011] Identify and remove the face region from the video to be processed to obtain the background region of the video to be processed;
[0012] The video to be processed is subjected to frame extraction to obtain the image set corresponding to the video to be processed;
[0013] The target network model is input into the video to be processed and the image set, and outputs a text description corresponding to the video to be processed.
[0014] In one optional implementation, the preset network model includes a video encoder, a text encoder, an image encoder, a multimodal joint encoder, a visual-text fusion network, a text decoder, and an image decoder, wherein:
[0015] The outputs of the video encoder, the text encoder, and the image encoder serve as the inputs of the multimodal joint encoder, and the output of the multimodal joint encoder serves as the input of the visual-text fusion network.
[0016] The text decoder and the image decoder are used to decode the output of the visual-text fusion network, and the text decoder is also used to output a text description.
[0017] In one optional implementation, training a preset network model using the video data to be trained includes:
[0018] The user video sample and its corresponding text description are input into the video encoder; the text description and text classification corresponding to the user video sample are input into the text encoder; the image set and the classification result of the user video sample based on the image set are input into the image encoder to train the preset network model until the preset network model converges.
[0019] In an optional implementation, before training a preset network model using the video data to be trained, the method further includes:
[0020] Perform text classification processing on the user video samples to obtain the text classification corresponding to the user video samples;
[0021] The user video sample is subjected to frame extraction processing to obtain the image set corresponding to the user video sample;
[0022] The image set is classified to obtain the classification result of the user video sample based on the image set.
[0023] In one optional implementation, the step of performing frame extraction processing on the user video samples to obtain the image set corresponding to the user video samples includes:
[0024] The user video sample is processed by extracting frames at a preset frame extraction rate to obtain multiple frames of images, which are then combined to form an image set corresponding to the user video sample.
[0025] In one optional implementation, the step of analyzing the text description corresponding to the video to be processed to determine the financial user profile corresponding to the video to be processed includes:
[0026] Extract keywords from the text description corresponding to the video to be processed;
[0027] Using a preset text model, based on the results of querying the keywords in the database of financial institutions, a description of the financial user profile corresponding to the video to be processed is generated.
[0028] Secondly, embodiments of this application provide a financial user profile prediction device, comprising:
[0029] The acquisition module is used to acquire training video data, which includes: user video samples of background areas extracted from a complete user video, text descriptions corresponding to the user video samples, text classifications of the user video samples, and an image set corresponding to the user video samples, wherein the user video samples are classified based on the image set.
[0030] The training module is used to train a preset network model using the video data to be trained, and obtain a target network model.
[0031] The processing module is used to process the video to be processed based on the target network model to obtain the text description corresponding to the video to be processed.
[0032] The analysis module is used to analyze the text description corresponding to the video to be processed to determine the financial user profile corresponding to the video to be processed.
[0033] Optionally, the processing module is specifically used for:
[0034] Identify and remove the face region from the video to be processed to obtain the background region of the video to be processed;
[0035] The video to be processed is subjected to frame extraction to obtain the image set corresponding to the video to be processed;
[0036] The target network model is input into the video to be processed and the image set, and outputs a text description corresponding to the video to be processed.
[0037] Optionally, the preset network model includes a video encoder, a text encoder, an image encoder, a multimodal joint encoder, a visual-text fusion network, a text decoder, and an image decoder, wherein:
[0038] The outputs of the video encoder, the text encoder, and the image encoder serve as the inputs of the multimodal joint encoder, and the output of the multimodal joint encoder serves as the input of the visual-text fusion network.
[0039] The text decoder and the image decoder are used to decode the output of the visual-text fusion network, and the text decoder is also used to output a text description.
[0040] Optionally, the training module is specifically used for:
[0041] The user video sample and its corresponding text description are input into the video encoder; the text description and text classification corresponding to the user video sample are input into the text encoder; the image set and the classification result of the user video sample based on the image set are input into the image encoder to train the preset network model until the preset network model converges.
[0042] Optionally, the financial user profiling prediction device further includes a preprocessing module for:
[0043] Before using the video data to be trained to train the preset network model, the user video samples are processed for text classification to obtain the text classification corresponding to the user video samples.
[0044] The user video sample is subjected to frame extraction processing to obtain the image set corresponding to the user video sample;
[0045] The image set is classified to obtain the classification result of the user video sample based on the image set.
[0046] Optionally, the preprocessing module is specifically used to perform frame extraction processing on the user video sample according to a preset frame extraction rate to obtain multiple frames of images, forming an image set corresponding to the user video sample.
[0047] Optionally, the analysis module is specifically used for:
[0048] Extract keywords from the text description corresponding to the video to be processed;
[0049] Using a preset text model, based on the results of querying the keywords in the database of financial institutions, a description of the financial user profile corresponding to the video to be processed is generated.
[0050] Thirdly, embodiments of this application also provide an electronic device, including a processor, an input device, an output device, and a memory, wherein the processor, input device, output device, and memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to invoke the program instructions to execute the method as described in the first aspect and any possible implementation thereof.
[0051] Fourthly, embodiments of this application provide a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the methods described in the first aspect and any possible implementation thereof.
[0052] This application embodiment acquires training video data, which includes: user video samples of background areas extracted from complete user videos, text descriptions corresponding to the user video samples, text classifications of the user video samples, and an image set corresponding to the user video samples, with the user video samples based on the classification results of the image set; a preset network model is trained using the training video data to obtain a target network model; the video to be processed is processed based on the target network model to obtain the text descriptions corresponding to the video to be processed; the text descriptions corresponding to the video to be processed are analyzed to determine the financial user profile corresponding to the video to be processed; in the scenario of user online videos, the model can be used to perform identification and analysis based on real-time user surroundings and background environment information, obtaining the corresponding description of the financial user profile while ensuring personal security information, which facilitates business personnel to timely and comprehensively judge the user's asset status and reasonably recommend products. Attached Figure Description
[0053] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0054] Figure 1 This is a flowchart illustrating a financial user profile prediction method provided in an embodiment of this application;
[0055] Figure 2 This is a schematic diagram of a model training and processing flow provided in an embodiment of this application;
[0056] Figure 3 This is a schematic diagram of the structure of a financial user profile prediction device provided in an embodiment of this application;
[0057] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0058] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. The term "embodiment" as used herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0059] Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0060] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0061] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0062] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0063] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0064] To better understand the embodiments of this application, the methods of applying the embodiments of this application will be described below.
[0065] The ImageNet project mentioned in this application embodiment is a large visualization database used for research on visual object recognition software.
[0066] The residual network mentioned in this embodiment is a convolutional neural network proposed by four researchers from Microsoft Research. Residual networks are characterized by their ease of optimization and the ability to improve accuracy by significantly increasing their depth. The residual blocks within them utilize skip connections, mitigating the vanishing gradient problem that arises from increasing depth in deep neural networks.
[0067] ChatGLM-6B, mentioned in the embodiments of this application, is an open-source dialogue language model that supports bilingual question-and-answer in Chinese and English, and has been optimized for Chinese.
[0068] The Transformer model mentioned in this application embodiment is a neural network that learns context and thus meaning by tracking relationships in sequence data (such as words in a sentence).
[0069] The Swin Transformer mentioned in this application embodiment is a Transformer-based deep learning model that has achieved state-of-the-art performance in visual tasks.
[0070] The contrastive learning mentioned in the embodiments of this application refers to aligning images and text into a joint feature space in a contrastive manner.
[0071] The multimodal fusion based on cross-attention mentioned in the embodiments of this application refers to fusing visual information into each layer of a language model with a cross-attention mechanism.
[0072] The terminal devices mentioned in the embodiments of this application include, but are not limited to, desktop computers and mobile terminals. The mobile terminals may include various handheld devices with wireless communication functions, computing devices such as laptops, or other processing devices connected to wireless modems, etc.
[0073] The method in this application aims to transcend the scope of ordinary face-to-face and video marketing. It primarily utilizes multimodal technology to model four types of information: video, text classification about the video, text descriptions based on the video, and image classification results related to the video. Building upon this, an image-text bimodal model is used to output text descriptions for images and videos containing customer background information. These descriptions not only help marketers read and understand customer background information more quickly but also help them match and compare customer profiles using keywords in the text. Considering privacy and information security concerns regarding facial information, this application avoids the facial region. Existing video cutout tools can be used to cut out the foreground portion containing faces, leaving only the background information, while using the customer's background information during face-to-face verification with marketers. The entire process avoids sensitive biometric information such as customer faces, ensuring privacy.
[0074] The method in this embodiment mainly involves using video data containing both the customer's background image and its background description as samples during training. Simultaneously, the model extracts text classification from the video, and then uses the model to classify the image set obtained by extracting frames from the video. During deployment, the bimodal model output of the image and text is processed by a text decoder to ultimately obtain a text description of the background of the video segment.
[0075] Please see Figure 1 This is a schematic flowchart illustrating a financial user profile prediction method provided in an embodiment of this application. Figure 1 As shown, the method may include:
[0076] 101. Obtain the training video data, which includes: user video samples of background regions extracted from the user's complete video, text descriptions corresponding to the user video samples, text classifications of the user video samples, image sets corresponding to the user video samples, and classification results of the user video samples based on the image sets.
[0077] The execution entity in this application embodiment can be a financial user profile prediction device, specifically, it can be implemented using electronic devices, which can be terminal devices.
[0078] This application primarily utilizes a pre-trained neural network model for video processing to achieve user profiling. First, the model and its training method are introduced. Detailed explanation of model training:
[0079] The aforementioned training video data is sample data used for model training, and mainly consists of four parts:
[0080] 1) User video samples are user videos that have undergone image cutout and anonymization processing, retaining only the background area;
[0081] 2) Text description for each video segment;
[0082] 3) Text classification for each video segment, such as "at home", "in the car", "playing with pets", etc.; specifically, it can be obtained by using a trained network model to perform text-based classification processing on the video. For example, the residual network ResNet50 can be used for text classification processing of the video. This application embodiment does not limit this.
[0083] 4) The image set obtained by extracting frames from the above video, and the classification result obtained by classifying the image set, such as "suit", "luxury house", "watch" etc.; specifically, the video can be processed by extracting frames according to a preset frame extraction rate to obtain the corresponding image set, and then a trained network model, such as the deep convolutional network DenseNet126, can be used to classify these images. This application embodiment does not limit this.
[0084] During online sales videos with customers, the camera captures the customer's image within the frame. Since facial features are highly sensitive biometric characteristics, using this information poses significant security and compliance risks. Therefore, in this embodiment, background / surrounding environment information from non-sensitive areas of the customer's face will be utilized. Specifically, existing video cutout tools will be used to extract the foreground portion containing the face, leaving only the background area.
[0085] 102. Use the above-mentioned video data to train the preset network model to obtain the target network model.
[0086] The preset network model used in this embodiment can learn features from both image and text modalities. The target network model obtained through training can process user videos and obtain textual descriptions of those videos, primarily descriptions related to the user's profile.
[0087] In one optional implementation, the aforementioned preset network model includes a video encoder, a text encoder, an image encoder, a multimodal joint encoder, a visual-text fusion network, a text decoder, and an image decoder, wherein:
[0088] The outputs of the aforementioned video encoder, text encoder, and image encoder serve as the inputs to the aforementioned multimodal co-encoder, and the output of the aforementioned multimodal co-encoder serves as the input to the aforementioned visual-text fusion network; the aforementioned text decoder and image decoder are used to decode the output of the aforementioned visual-text fusion network, and the aforementioned text decoder is also used to output text descriptions.
[0089] Specifically, the pre-defined network model processes video data through a video encoder, text data through a text encoder, and image data through an image encoder. Then, it performs fusion processing based on a multimodal joint encoder. The visual-text fusion network learns the mapping relationship between image and text modalities in the feature space. Finally, the output of the visual-text fusion network enters the text decoder, which can output a text description of the video.
[0090] The visual-text fusion network mentioned in this application embodiment is a bimodal model that can process text and images simultaneously, also known as a joint visual-language model, such as OpenAI's CLIP (Contrastive Language-Image Pre-Training) model.
[0091] Further, optionally, the above-mentioned training of a preset network model using the video data to be trained includes:
[0092] The user video samples and their corresponding text descriptions are input into the video encoder; the text descriptions and text classifications corresponding to the user video samples are input into the text encoder; the image set and the classification results of the user video samples based on the image set are input into the image encoder to train the preset network model until the preset network model converges.
[0093] Specifically, the user's video sample (containing only the background) and the text description of the video segment are fed into the video encoder, such as a video encoder built with VisionTransformer;
[0094] Meanwhile, the text description and text classification of the video segment are fed into a text encoder, for example, a text encoder built with chatGLM-6B can be used;
[0095] The image set obtained from frame extraction, along with its own text description, is fed into the image encoder. For example, an image encoder built using SwingTransformer can be used.
[0096] The video encoder, text encoder, and image encoder described above output information from their respective modalities, which is then input into the multimodal joint encoder.
[0097] The multimodal joint encoder in this embodiment can consist of three Transformer encoders, each containing a self-attention module and a feedforward module. There is a cross-attention mechanism between the three Transformers. Furthermore, all feedforward networks and self-attention modules share weights.
[0098] The output of the multimodal co-encoder is a vector in a high-dimensional space, which will be fed into the visual-text fusion network. A dual-modal model, VisualTextDualEncoder, can be used. This network is constructed using the CLIP framework based on contrastive learning, which can learn excellent cross-modal joint representation capabilities. Through the contrastive learning mechanism, a well-trained CLIP model can learn the mapping relationship between image and text modalities in the feature space.
[0099] Since the user's background should be observed globally, this transformer's global attention can effectively capture the correlation between different parts of the background, thereby extracting important deep features from the surrounding environment behind the customer.
[0100] Based on the learned information, the CLIP model can use only the intermediate text generated during inference as an image caption model, and output the corresponding text description.
[0101] In an optional implementation, before training a preset network model using the aforementioned video data to be trained, the method further includes:
[0102] The above user video samples are subjected to text classification processing to obtain the text classification corresponding to the above user video samples;
[0103] Frame extraction is performed on the above user video samples to obtain the image set corresponding to the above user video samples;
[0104] The above image set is classified to obtain the classification results of the above user video samples based on the above image set.
[0105] Further, optionally, the above-mentioned frame extraction processing of the user video samples to obtain the image set corresponding to the user video samples includes:
[0106] The aforementioned user video samples are processed by frame extraction at a preset frame extraction rate to obtain multiple frames of images, which together form the image set corresponding to the aforementioned user video samples.
[0107] Specifically, the preprocessing of the video data to be trained has already been mentioned above, including text classification, frame extraction, and image classification. In one implementation, the following processing can be performed:
[0108] The ResNet50 residual network, pre-trained on ImageNet, is used to perform text-based classification on the entire video to obtain the text classification of the video.
[0109] The preset frame rate can be set as needed, for example, 0.5 FPS. This involves fetching frames from the video every 2 seconds to obtain a set of images. For instance, a 100-second video of a customer talking to a salesperson in their car would ultimately yield a set of 50 images.
[0110] The DenseNet126 deep convolutional network, pre-trained on ImageNet, is used to classify the image set, yielding classification results based on the image set. It's important to note that this stores the classification results, not the images themselves containing facial information.
[0111] In one specific implementation, model training data can be selected from existing business tables and divided into training, validation, and test sets. For example, approximately 10MB of customer data can be selected from existing business tables as a dataset, and the data can be randomly divided into training, validation, and test sets in an 8:1:1 ratio. All three datasets are balanced, meaning that data belonging to the same customer can only appear in one of the training, validation, or test sets simultaneously.
[0112] Model training can employ gradient descent optimization based on the AdamLR algorithm, with cross-entropy used as the loss optimization function during the training process. The final prediction result is based on the results on the test set. The learning strategy can be set or adjusted as needed. For example, the initial learning rate can be set to e-3, reduced to e-4 after 10 warm-up steps, and then reduced by 1 / 10 every 20 steps. This embodiment does not impose any limitations on this.
[0113] Once the target network model has been trained, it can be put into application, i.e., step 103 can be executed.
[0114] 103. Based on the above target network model, process the video to be processed to obtain the text description corresponding to the video to be processed.
[0115] The aforementioned videos to be processed are also user videos where only the background area is retained.
[0116] Specifically, in practical applications, the above model can be used in various video scenarios of financial services to process user videos in real time, helping business personnel to identify and analyze user profiles in a timely manner, such as when customers communicate with business personnel via video to introduce financial products.
[0117] In an optional implementation, step 103 includes:
[0118] 31. Identify and remove the face regions from the above-mentioned video to be processed, and obtain the background region of the video to be processed;
[0119] 32. Perform frame extraction on the above-mentioned video to be processed to obtain the image set corresponding to the above-mentioned video to be processed;
[0120] 33. Input the above-mentioned video to be processed and the above-mentioned image set into the above-mentioned target network model, and output the text description corresponding to the above-mentioned video to be processed.
[0121] Similar to model training, the input video to the model is a video with the user's face region removed to ensure user privacy. The above steps can be found in the detailed description of the relevant steps during model training, and will not be repeated here.
[0122] The above frame extraction process is the same as during model training. The preset frame extraction rate can be set as needed, such as 0.5 FPS, which means extracting a frame every 2 seconds to obtain an image set.
[0123] Figure 2 This is a schematic diagram illustrating a model training and processing flow provided in an embodiment of this application. The model structure and model training flow involved in the foregoing embodiments can also be referred to... Figure 2 As shown. Specifically, when using a target network model to analyze the financial information of video users and obtain predicted text descriptions, the process can be as follows:
[0124] The user's video is acquired, and the background information other than faces and foreground is extracted from the user's video by the preprocessing video tool, which is the video to be processed;
[0125] The video to be processed is subjected to frame extraction at 0.5 FPS to obtain the corresponding image set;
[0126] The sets of videos and images to be processed are fed into the pre-trained video encoder and image encoder, respectively.
[0127] Then, the outputs of the two encoders are fed into the multimodal co-encoder, and then into the CLIP-based image-text dual-modal model VisualTextDualEncoder;
[0128] The output of the bimodal model is fed into the text decoder, which ultimately outputs a text description of the video.
[0129] In one implementation, the model's output can be a text description, such as:
[0130] "A person dressed in xxx is in an office of xxx, where there is an expensive xxx painting on the wall, the decor is grand and classy, and there is a luxurious handicraft on the bookshelf behind them."
[0131] It is important to note that the user's clothing information here is derived from the classification results of the foreground image encoder, rather than from the image itself containing facial information.
[0132] 104. Analyze the text descriptions corresponding to the videos to be processed to determine the financial user profiles corresponding to the videos to be processed.
[0133] Furthermore, the text descriptions output by the model can be used to make preliminary judgments about user profiles, and to help determine information such as the user's asset status and consumption habits.
[0134] In an optional implementation, step 104 includes:
[0135] Extract keywords from the text descriptions corresponding to the videos to be processed;
[0136] Using a pre-defined text model, based on the results of querying the aforementioned keywords in the databases of financial institutions, a description of the financial user profile corresponding to the aforementioned video to be processed is generated.
[0137] Specifically, a text model based on chatGLM-6B can be used to perform keyword comparison and query the financial institution's database to define the user's profile. Simultaneously, business personnel can also directly assess the customer's asset situation and other information based on this text description.
[0138] This application primarily proposes a method for assessing a customer's asset status in an online video conversation scenario, based on real-time information about the customer's surroundings and background (video), thereby enabling financial institutions to identify high-net-worth clients during sales activities. Compared to traditional methods, this approach offers the following significant advantages:
[0139] 1) Security. Throughout the entire process, no facial features (highly sensitive information) of the customer are extracted, thus eliminating risks such as personal information leakage.
[0140] 2) Highly complementary to traditional methods. Unstructured video and image data are difficult to efficiently extract effective customer information. The method in this application utilizes multimodal artificial intelligence technology to successfully textualize visual information, making user profile matching and high-net-worth customer discovery easier.
[0141] 3) Each of the three modalities uses its own encoder, which can minimize information loss in the low-dimensional space.
[0142] 4) Multimodal joint encoders can fuse information sources from different modalities at low dimensions by sharing weights and cross-attention, thereby reducing the gap between modalities.
[0143] 5) The CLIP framework, which uses contrastive learning, can learn the relationships between text and image modalities to the maximum extent and output the most suitable text description.
[0144] 6) The entire process is highly automated, which greatly reduces the workload and time of operators, and allows business personnel to master the information of unstructured data to the greatest extent possible, laying the groundwork for their subsequent high-net-worth client mining.
[0145] For different application scenarios or needs, the model structure, training method or training data can be adaptively adjusted, and the embodiments of this application do not limit this.
[0146] Based on the aforementioned method embodiments, this application also provides a financial user profile prediction device.
[0147] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a financial user profiling prediction device provided in an embodiment of this application. Figure 3 As shown, the financial user profiling prediction device 300 includes:
[0148] The acquisition module 310 is used to acquire training video data, which includes: user video samples of background regions extracted from the user's complete video, text descriptions corresponding to the user video samples, text classifications of the user video samples, and image sets corresponding to the user video samples, with the user video samples being classified based on the image sets.
[0149] Training module 320 is used to train a preset network model using the above-mentioned video data to be trained, and obtain the target network model;
[0150] Processing module 330 is used to process the video to be processed based on the above target network model to obtain the text description corresponding to the video to be processed.
[0151] The analysis module 340 is used to analyze the text description corresponding to the video to be processed to determine the financial user profile corresponding to the video to be processed.
[0152] Optionally, the aforementioned processing module 330 is specifically used for:
[0153] Identify and remove the face regions from the video to be processed to obtain the background region of the video to be processed;
[0154] The above-mentioned video to be processed is subjected to frame extraction to obtain the image set corresponding to the above-mentioned video to be processed;
[0155] The above-mentioned video to be processed and the above-mentioned image set are input into the above-mentioned target network model, and the text description corresponding to the above-mentioned video to be processed is output.
[0156] Optionally, the aforementioned preset network model includes a video encoder, a text encoder, an image encoder, a multimodal joint encoder, a visual-text fusion network, a text decoder, and an image decoder, wherein:
[0157] The outputs of the aforementioned video encoder, text encoder, and image encoder serve as the inputs of the aforementioned multimodal joint encoder, and the output of the aforementioned multimodal joint encoder serves as the inputs of the aforementioned visual-text fusion network.
[0158] The aforementioned text decoder and image decoder are used to decode the output of the aforementioned visual-text fusion network. The aforementioned text decoder is also used to output text descriptions.
[0159] Optionally, the above training module 320 is specifically used for:
[0160] The user video samples and their corresponding text descriptions are input into the video encoder; the text descriptions and text classifications corresponding to the user video samples are input into the text encoder; the image set and the classification results of the user video samples based on the image set are input into the image encoder to train the preset network model until the preset network model converges.
[0161] Optionally, the aforementioned financial user profiling prediction device further includes a preprocessing module 350, used for:
[0162] Before training the preset network model using the above-mentioned video data to be trained, the above-mentioned user video samples are subjected to text classification processing to obtain the text classification corresponding to the above-mentioned user video samples.
[0163] Frame extraction is performed on the above user video samples to obtain the image set corresponding to the above user video samples;
[0164] The above image set is classified to obtain the classification results of the above user video samples based on the above image set.
[0165] Optionally, the preprocessing module 350 is specifically used to perform frame extraction processing on the user video sample according to a preset frame extraction rate to obtain multiple frames of images and form an image set corresponding to the user video sample.
[0166] Optionally, the aforementioned analysis module 340 is specifically used for:
[0167] Extract keywords from the text descriptions corresponding to the videos to be processed;
[0168] Using a pre-defined text model, based on the results of querying the aforementioned keywords in the databases of financial institutions, a description of the financial user profile corresponding to the aforementioned video to be processed is generated.
[0169] According to the specific implementation of the embodiments of this application, Figure 1 and Figure 2 The steps involved in the financial user profiling prediction method shown can be from... Figure 3 The various modules in the financial user profiling prediction device 300 shown are responsible for this, and will not be described in detail here.
[0170] The financial user profile prediction device 300 of this application embodiment acquires training video data, which includes: user video samples of background areas extracted from a complete user video, text descriptions corresponding to the user video samples, text classifications of the user video samples, and an image set corresponding to the user video samples, wherein the user video samples are classified based on the image set; a preset network model is trained using the training video data to obtain a target network model; the video to be processed is processed based on the target network model to obtain text descriptions corresponding to the video to be processed; and the text descriptions corresponding to the video to be processed are analyzed to determine the financial user profile corresponding to the video to be processed. In scenarios involving online user videos, the model can perform identification and analysis based on real-time user surroundings and background environment information, ensuring personal security information while obtaining corresponding descriptions of financial user profiles, facilitating business personnel to timely and comprehensively assess user asset status and reasonably recommend products.
[0171] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. For example... Figure 4 As shown, the electronic device 400 includes a processor 401 and a memory 402. The electronic device 400 may also include a bus 403, through which the processor 401 and the memory 402 can be interconnected. The bus 403 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 403 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 4 The bus is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The electronic device 400 may further include an input / output device 404, which may include a display screen, such as a liquid crystal display (LCD). The memory 402 is used to store one or more programs containing instructions; the processor 401 is used to invoke the instructions stored in the memory 402 to execute the above-mentioned... Figure 1 or Figure 2Some or all of the method steps mentioned in the embodiments.
[0172] It should be understood that, in the embodiments of this application, the processor 401 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0173] Input device 402 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 403 may include a display (LCD, etc.), a speaker, etc.
[0174] The memory 404 may include read-only memory and random access memory, and provides instructions and data to the processor 401. A portion of the memory 404 may also include non-volatile random access memory. For example, the memory 404 may also store device type information.
[0175] Through the electronic device 400 of this application embodiment, the electronic device 400 can acquire training video data, which includes: user video samples of background areas extracted from a complete user video, text descriptions corresponding to the user video samples, text classifications of the user video samples, and an image set corresponding to the user video samples, wherein the user video samples are classified based on the image set; a preset network model is trained using the training video data to obtain a target network model; the video to be processed is processed based on the target network model to obtain the text descriptions corresponding to the video to be processed; the text descriptions corresponding to the video to be processed are analyzed to determine the financial user profile corresponding to the video to be processed; in the scenario of user online video, the model can be used to perform identification and analysis based on real-time user surrounding and background environment information, ensuring personal security information while obtaining the corresponding description of the financial user profile, which facilitates business personnel to timely and comprehensively judge the user's asset status and reasonably recommend products.
[0176] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the financial user profiling prediction methods described in the above method embodiments.
[0177] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0178] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical or other forms.
[0179] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0180] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
Claims
1. A financial user profiling prediction method, characterized in that, The method includes: Acquire training video data, which includes: user video samples of background regions extracted from complete user videos, text descriptions corresponding to the user video samples, text classifications of the user video samples, and image sets corresponding to the user video samples, with the user video samples being classified based on the image sets. Before training a preset network model using the video data to be trained, the method further includes: The user video sample is subjected to text classification processing to obtain the text classification corresponding to the user video sample; the user video sample is subjected to frame extraction processing to obtain the image set corresponding to the user video sample; the image set is subjected to classification processing to obtain the classification result of the user video sample based on the image set; The preset network model is trained using the video data to be trained to obtain the target network model; The preset network model includes a video encoder, a text encoder, an image encoder, a multimodal joint encoder, a visual-text fusion network, a text decoder, and an image decoder, wherein: the outputs of the video encoder, the text encoder, and the image encoder serve as the inputs of the multimodal joint encoder, and the output of the multimodal joint encoder serves as the input of the visual-text fusion network; the text decoder and the image decoder are used to decode the output of the visual-text fusion network, and the text decoder is also used to output text descriptions; The step of training a preset network model using the video data to be trained specifically includes: inputting the user video sample and the corresponding text description into the video encoder; inputting the corresponding text description and text classification into the text encoder; inputting the image set and the classification result of the user video sample based on the image set into the image encoder, and training the preset network model until the preset network model converges; The target network model is used to process the video to be processed to obtain the text description corresponding to the video to be processed. The analysis of the text description corresponding to the video to be processed determines the financial user profile corresponding to the video to be processed. Specifically, this includes: extracting keywords from the text description corresponding to the video to be processed; and using a preset text model, generating a description of the financial user profile corresponding to the video to be processed based on the results of querying the keywords in the database of financial institutions.
2. The method of claim 1, wherein, The step of performing frame extraction processing on the user video samples to obtain the image set corresponding to the user video samples includes: The user video sample is processed by extracting frames at a preset frame extraction rate to obtain multiple frames of images, which are then combined to form an image set corresponding to the user video sample.
3. A financial user profiling prediction apparatus characterized by, For performing the method as described in claim 1 or 2, the apparatus comprises: The acquisition module is used to acquire training video data, which includes: user video samples of background areas extracted from a complete user video, text descriptions corresponding to the user video samples, text classifications of the user video samples, and an image set corresponding to the user video samples, wherein the user video samples are classified based on the image set. The training module is used to train a preset network model using the video data to be trained, and obtain a target network model. The processing module is used to process the video to be processed based on the target network model to obtain the text description corresponding to the video to be processed. The analysis module is used to analyze the text description corresponding to the video to be processed to determine the financial user profile corresponding to the video to be processed.
4. An electronic device, comprising: The system includes a processor, an input device, an output device, and a memory, which are interconnected. The memory is used to store a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute the method as described in claim 1 or 2.
5. A computer storage medium, characterized in that The computer storage medium stores a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the method as described in any one of claims 1 or 2.