Business processing method, device, equipment, medium and product

By acquiring and analyzing the voice and video data of target users, a business knowledge graph is generated, which solves the problem of low efficiency and accuracy of terminal devices in identifying business needs in complex environments, and realizes efficient business processing through multimodal fusion.

CN121787535APending Publication Date: 2026-04-03INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing terminal devices struggle to effectively identify target users' business needs in complex deployment environments. The combination of voice recognition and lip-reading technologies has limited effectiveness, resulting in low recognition efficiency and accuracy, which impacts customer service quality.

Method used

By acquiring the target user's voice and video data, we can determine the user's voice text, voiceprint features, and lip movement features, generate a business knowledge graph, and combine these features to generate request processing strategies and provide feedback on the processing results.

Benefits of technology

It achieves multimodal fusion of biometric recognition and visual semantic understanding, which improves the efficiency and accuracy of business requirement identification, enhances the accuracy and reliability of business processing results, and ensures service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787535A_ABST
    Figure CN121787535A_ABST
Patent Text Reader

Abstract

The invention discloses a business processing method, device and equipment, a medium and a product, which are applied to the technical field of finance. The method comprises the following steps: in response to a service processing request initiated by a target user, obtaining user voice data and user video data of the target user; determining a user voice text and a user voiceprint feature of the target user according to the user voice data, and determining a lip motion feature of the target user according to the user video data; generating a business knowledge graph corresponding to the target user according to the user voice text, the user voiceprint features and the lip motion features; and according to the service knowledge graph, generating a request processing strategy for the target user, and according to the request processing strategy, generating a processing result of a service request initiated by the target user and feeding back the processing result. According to the embodiment, the business knowledge graph of the target user can be generated, the request processing result and feedback of the business request of the target user can be obtained, and the identification efficiency and accuracy of the business requirement of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial technology, and in particular to a business processing method, apparatus, equipment, medium, and product. Background Technology

[0002] With the continuous improvement of banking technology and the development of terminal devices, obtaining voice and video data of the target user when the target user initiates a business request through terminal devices, generating a request processing strategy for the initiated business request, obtaining the request processing result and providing feedback has become the main form of customer service.

[0003] However, as business requests become increasingly complex and the deployment environments of terminal devices become more intricate, existing methods for handling these requests struggle to cope with the growing complexity. Specifically, traditional speech recognition technologies have limited effectiveness in recognizing target users' speech data in non-steady-state noise environments, leading to a sharp drop in the efficiency and accuracy of identifying target users' business needs. Meanwhile, lip-syncing technology, as a visual aid, can infer speech content by acquiring the target user's lip movement features, providing multimodal data support to enhance speech recognition results. However, existing lip-syncing technologies offer extremely limited enhancement to speech recognition results and struggle to deeply integrate the recognition results with the target user's speech data, resulting in low accuracy in identifying user needs and further compromising the quality of customer service.

[0004] Therefore, how to improve the efficiency and accuracy of identifying the business needs of target users, enhance the accuracy and reliability of business processing results, and ensure the service quality for target users has become an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] This invention provides a business processing method, apparatus, equipment, medium, and product to improve the efficiency and accuracy of identifying the business needs of target users, enhance the accuracy and reliability of business processing results, and ensure the quality of service to target users.

[0006] According to one aspect of the present invention, a business processing method is provided, comprising:

[0007] In response to a business processing request initiated by a target user, the user's voice data and user video data are obtained;

[0008] Based on the user voice data, determine the user voice text and user voiceprint features of the target user, and based on the user video data, determine the lip movement features of the target user;

[0009] Based on the user's voice text, the user's voiceprint features, and the lip movement features, a business knowledge graph corresponding to the target user is generated;

[0010] Based on the business knowledge graph, a request processing strategy for the target user is generated, and the processing result of the business request initiated by the target user is generated and fed back based on the request processing strategy.

[0011] According to another aspect of the present invention, a business processing apparatus is provided, comprising:

[0012] The data acquisition module is used to acquire the user's voice data and user video data in response to a business processing request initiated by the target user.

[0013] The data parsing module is used to determine the user's voice text and user voiceprint features of the target user based on the user's voice data, and to determine the lip movement features of the target user based on the user's video data.

[0014] The knowledge graph generation module is used to generate a business knowledge graph corresponding to the target user based on the user's voice text, the user's voiceprint features, and the lip movement features.

[0015] The request processing module is used to generate a request processing strategy for the target user based on the business knowledge graph, and generate and feed back the processing result of the business request initiated by the target user based on the request processing strategy.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the business processing method described in any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the business processing method described in any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the business processing method described in any embodiment of the present invention.

[0021] The technical solution of this invention, in response to a business processing request initiated by a target user, acquires the target user's voice data and video data; determines the target user's voice text and voiceprint features based on the voice data, and determines the target user's lip movement features based on the video data; generates a business knowledge graph corresponding to the target user based on the voice text, voiceprint features, and lip movement features; generates a request processing strategy for the target user based on the business knowledge graph, and generates and feeds back the processing result of the business request initiated by the target user based on the request processing strategy. This implementation scheme can determine the target user's voice text, voiceprint features, and lip movement features based on the target user's voice and video data, and generate a business knowledge graph for the target user accordingly, thereby generating a request processing strategy for the target user's business requests and feeding back the request processing result. This achieves multimodal fusion of biometric recognition and visual semantic understanding, improves the efficiency and accuracy of identifying the target user's business needs, enhances the accuracy and reliability of business processing results, and ensures the service quality for the target user.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a business processing method provided according to Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart of a business processing method provided according to Embodiment 2 of the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of a business processing device according to Embodiment 3 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the business processing method of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This invention provides a flowchart of a business processing method according to Embodiment 1. This embodiment is applicable to situations where the identification efficiency and accuracy of the target user's business needs are low during the processing of business requests. This method can be executed by a business processing device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0032] S110. In response to a business processing request initiated by the target user, obtain the target user's voice data and video data.

[0033] User voice data refers to audio information generated when a target user uses voice interaction functions. User video data refers to image information generated when a target user uses video-related services, specifically image / video information of the target user's lip area. Specifically, when a target user processes a business request through a self-service terminal device—such as a balance inquiry, loan application inquiry, or financial product consultation request—user voice data can be collected using a voice acquisition device pre-deployed in the self-service terminal device, such as a microphone array. Similarly, user video data can be collected using an image acquisition device pre-deployed in the self-service terminal device, such as a camera.

[0034] Furthermore, after acquiring the user's voice data and user video data, at least one audio sampling frame in the user's voice data and a first timestamp corresponding to each audio sampling frame can be determined, as well as at least one video image frame in the user's video data and a second timestamp corresponding to each video image frame. Based on the time bases corresponding to the user's voice data and user video data respectively, such as the time base corresponding to the user's voice data being 48kHz and the time base corresponding to the user's video data being 90kHz, the first timestamp and the second timestamp are unified into the same clock domain. For example, the time base corresponding to the user's voice data being 48kHz can be converted to 90kHz to obtain the audio sampling frames corresponding to each video image frame, ensuring audio-visual synchronization between the voice signals and video signals in the acquired user's voice data and user video data.

[0035] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant regions.

[0036] S120. Based on the user's voice data, determine the user's voice text and user voiceprint features of the target user, and based on the user's video data, determine the lip movement features of the target user.

[0037] The user's voice text can be the text content obtained by recognizing the user's voice data using speech recognition technology. The user's voiceprint features can be biometric features in the user's voice data used to uniquely identify the target user, specifically including pitch, timbre, and speech rate. Specifically, based on the target user's voice data and a pre-defined speech recognition algorithm, speech recognition can be performed to determine the text content within the target user's voice data. Simultaneously, at least one speech acoustic feature in the target user's voice data can be analyzed and statistically analyzed, including the spectrum, fundamental frequency, and formants, to obtain the target user's speech feature parameters. Based on these speech feature parameters, the target user's voiceprint features can be determined.

[0038] Among them, lip movement features can be the dynamic changes in the lip region of the target user during voice interaction. Specifically, each video image frame in the user's video can be input into a pre-trained neural network model to obtain at least one feature parameter in each video image frame used to quantify the lip movement of the target user. Based on each feature parameter, the lip movement features of the target user can be determined, where the feature parameters can include geometric features, texture features, and optical flow features.

[0039] Optionally, based on the user's voice data, the user's voice text and user voiceprint features are determined, including: based on the user's voice data and a preset voice separation algorithm, the target user's voice data and user voiceprint features are determined; and the target voice data is subjected to voice recognition to obtain the target user's voice text.

[0040] The user voice data can include target voice data and environmental noise data. The target voice data can be audio data describing a business processing request initiated by the target user, such as a request to query transaction records and account balance in a bank card. Specifically, a preset voice separation algorithm can be used to separate the user voice data into target voice data and environmental noise data. This voice separation algorithm can utilize the spatial information corresponding to multiple microphones, using beamforming to enhance the target voice data and suppress environmental noise data, thus obtaining the target voice data. Alternatively, it can determine the mixed speech spectrum corresponding to the user voice data, and use a clustering algorithm to classify the spectral units in the mixed speech spectrum to obtain the target voice data. Then, voice feature analysis can be performed on the target voice data to obtain at least one voice acoustic feature. Based on these voice acoustic features, the target user's voice feature parameters are determined, and based on these voice feature parameters, the target user's voiceprint features are determined.

[0041] Furthermore, to obtain the target user's speech text by performing speech recognition on the target speech data, specifically, at least one speech feature parameter in the target speech data can be determined by short-time Fourier transform. Based on each speech feature parameter and a preset acoustic model, the speech factors associated with each speech feature parameter can be determined, and the target user's speech text can be obtained through these speech factors.

[0042] This technical solution can determine the target user's voice text and voiceprint features based on the target user's voice data, thereby increasing the recognition of the target user's biometric features during business processing and improving the efficiency and accuracy of recognizing the target user's business needs.

[0043] Optionally, based on user video data, the lip movement features of the target user are determined, including: generating at least one lip image video frame of the target user based on user video data; extracting features from multiple lip feature key points associated with each lip image video frame and obtaining the motion trajectory corresponding to each lip feature key point; and determining the lip movement features of the target user based on the motion trajectory corresponding to each lip feature key point.

[0044] Specifically, video processing tools can extract at least one video image frame from the user's video data at preset time intervals, such as 30 frames per second. Computer vision techniques, such as a lip feature point detection model, are then used to detect at least one lip image video frame of the target user in each video image frame. Next, feature extraction is performed on multiple lip feature key points associated with each lip image video frame. These lip feature key points can include the corners of the lips, the cupid's bow, and the lip edge. Feature vectors for each lip feature key point are calculated across multiple consecutive lip image video frames to obtain the motion trajectory corresponding to each lip feature key point. Furthermore, motion feature parameters for each lip feature key point can be determined based on its corresponding motion trajectory. Based on these motion feature parameters, the lip motion characteristics of the target user are determined. These motion feature parameters can include the curvature, inflection points, period, and frequency of the movement trajectory of each lip feature key point during the target user's voice interaction.

[0045] This technical solution can determine the lip movement features of target users based on their video data. The solution incorporates multimodal fusion of visual semantic understanding, which improves the efficiency and accuracy of identifying the business needs of target users and enhances the accuracy and reliability of business processing results.

[0046] S130. Generate a business knowledge graph corresponding to the target user based on the user's voice text, user voiceprint features, and lip movement features.

[0047] The business knowledge graph can be a tool for displaying and managing various characteristic information related to target users and their interrelationships. This characteristic information can include the target user's user information, speech-text features, voiceprint features, and lip movement features. User information can include the user's name, age, occupation, and financial account data. Specifically, for example, a target user might express their intention to purchase financial products through voice interaction with a self-service terminal. During this voice interaction, the self-service terminal can not only analyze the speech-text content of the target user's voice data but also extract the user's voiceprint features, such as a calm demeanor, moderate speaking speed, and lip movement features, to help confirm the accuracy of speech recognition and determine the clarity and standard of the user's pronunciation.

[0048] For example, during voice interaction, the following characteristic information related to the target user can be obtained: Zhang San, male, 35 years old, occupation: teacher; frequently uses financial terminology, indicating some understanding of financial business; calm demeanor, clear intention to purchase financial products; bank account balance: 100,000 yuan, account opened for five years. Based on the obtained characteristic information, the following relationships can be obtained: the target user has a bank account with a balance of 100,000 yuan and has been open for five years; the target user is interested in financial products; and the target user expresses stable emotions and clear intentions. Furthermore, based on the obtained user characteristic information of the target user and the relationships between these user characteristics, a business knowledge graph corresponding to the target user can be generated. This embodiment does not impose specific limitations on this.

[0049] S140. Based on the business knowledge graph, generate a request processing strategy for the target user, and generate and feed back the processing results of the business requests initiated by the target user based on the request processing strategy.

[0050] The request processing strategy can be the request processing methods and processes executed for business requests initiated by a target user. Specifically, based on the business knowledge graph, the target user's business needs and user profile can be determined. Based on the target user's business needs and user profile, a request processing strategy for the target user is generated. Then, according to the request processing strategy, the business requests initiated by the target user are processed, and the processing results are obtained and fed back to the target user.

[0051] For example, continuing the previous example, the specific strategy for processing the target user's request could be as follows: Based on the target user's characteristic information and its correlations, if the target user's voice and text characteristics indicate a desire to purchase financial products, then the user's needs can be pinpointed to how to purchase financial products and inquire about related product details. Based on user information, such as the target user's age being 35 and occupation being a teacher, it can be determined that the target user's risk tolerance is moderate. Based on this, financial products with moderate risk and optimal returns can be recommended to the target user. Furthermore, based on voiceprint and lip movement characteristics, if it is determined that the target user's emotions are stable and their intentions are clear, an automated process can be used to guide the user to inquire about product details, complete a risk assessment, and assist the target user in completing the purchase process. Further, the automated process can continuously monitor the target user's emotions. Specifically, based on the target user's voiceprint and lip movement characteristics, the target user's emotions can be monitored in real time. If the user's emotions change from stable to impatient and agitated, then a human customer service representative can be contacted to explain the details of the financial product and guide the user to complete the purchase. This embodiment does not impose specific limitations on this.

[0052] Optionally, after generating a request processing strategy for the target user based on the business knowledge graph, and generating and feeding back the processing result of the business request initiated by the target user based on the request processing strategy, the method further includes: determining whether the target user's request processing strategy meets the pre-set request processing strategy optimization requirements based on the target user's user voice text, user voiceprint features, and lip movement features; if so, adjusting the request processing strategy and obtaining the adjusted request processing strategy.

[0053] Specifically, the target user's emotions can be monitored in real time based on the text content of the user's voice-text, as well as the user's voiceprint and lip movement characteristics. These emotions can include positive, negative, neutral, angry, and anxious. Furthermore, based on the real-time monitoring results, it can be determined whether the request processing strategy for the target user meets the pre-set optimization requirements. For example, if the target user's voiceprint shows a pleasant mood at the end of the service process, their lip movement shows satisfaction, and their voice-text provides a positive evaluation of customer service, then the request processing strategy for the target user is normal and does not meet the pre-set optimization requirements. Conversely, if the user's voice-text contains the following message: "Your service is terrible, I've been waiting so long and it's still not resolved," and the voiceprint and lip movement characteristics indicate that the target user is anxious and their voice volume is significantly increased, then the target user is exhibiting anger. If the request processing strategy for the target user is found to be abnormal and meets the pre-set optimization requirements, the request processing strategy can be adjusted. Specifically, the automated process can be switched to a manual service, whereby a person will handle the business requests of the target user and simplify the business request processing flow.

[0054] This technical solution can combine the target user's voice text, voiceprint features, and lip movement features to adjust the generated request processing strategy, thereby improving the accuracy and reliability of business processing results and ensuring the service quality for the target user.

[0055] The technical solution of this invention, in response to a business processing request initiated by a target user, acquires the target user's voice data and video data; determines the target user's voice text and voiceprint features based on the voice data, and determines the target user's lip movement features based on the video data; generates a business knowledge graph corresponding to the target user based on the voice text, voiceprint features, and lip movement features; generates a request processing strategy for the target user based on the business knowledge graph, and generates and feeds back the processing result of the business request initiated by the target user based on the request processing strategy. This implementation scheme can determine the target user's voice text, voiceprint features, and lip movement features based on the target user's voice and video data, and generate a business knowledge graph for the target user accordingly, thereby generating a request processing strategy for the target user's business requests and feeding back the request processing result. This achieves multimodal fusion of biometric recognition and visual semantic understanding, improves the efficiency and accuracy of identifying the target user's business needs, enhances the accuracy and reliability of business processing results, and ensures the service quality for the target user.

[0056] Example 2

[0057] Figure 2 This is a flowchart of a business processing method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further optimizes the above business processing method.

[0058] Furthermore, the step "generating a business knowledge graph corresponding to the target user based on user voice text, user voiceprint features, and lip movement features" is refined to "identifying the text content of the user's voice text to obtain business content keywords; determining the target user's user behavior profile based on user voiceprint features and lip movement features; and generating a business knowledge graph corresponding to the target user based on the business content keywords and user behavior profile." This improves the processing method for business requests initiated by the target user. Figure 2 As shown, the method includes:

[0059] S210. In response to a business processing request initiated by the target user, obtain the target user's voice data and video data.

[0060] S220. Based on the user's voice data, determine the user's voice text and user voiceprint features of the target user, and based on the user's video data, determine the lip movement features of the target user.

[0061] S230. Recognize the text content of the user's voice text to obtain business content keywords.

[0062] The business content keywords can be content information used to describe and identify the business needs in a business request initiated by a target user. For example, if the text content of a user's voice message is "Recommend a stable wealth management product with a low minimum investment amount, and tell me what the annualized rate of return is" or "Can I redeem the net asset value wealth management product I bought before, and how is the return calculated?", then the corresponding business content keywords obtained from the voice message are: annualized rate of return, net asset value wealth management, stable wealth management, minimum investment amount, and redemption. This embodiment does not impose specific limitations on this.

[0063] S240. Based on the user's voiceprint characteristics and lip movement characteristics, determine the user behavior profile of the target user.

[0064] The user behavior profile can be a set of data identifiers used to identify the target user. For example, the acquired voiceprint characteristics of the target user may include: fast speech rate, high-pitched and highly variable pitch, and high volume; the lip movement characteristics of the target user may include: large amplitude, high frequency, and stiff lip posture. Based on the user's voiceprint and lip movement characteristics, the target user's behavior profile can be determined as: a young, impatient credit user with immediate business needs, emotional anxiety, and a focus on installment efficiency and cost; that is, the target user is identified as a high-frequency credit card user. This embodiment does not impose specific limitations on this.

[0065] Based on the user's voiceprint features and lip movement features, determine the target user's user behavior profile, including: determining the target user's basic user features based on the user's voiceprint features, and determining the target user's user behavior features based on the lip movement features; and determining the target user's user behavior profile based on the user's basic user features and user behavior features.

[0066] The user's basic characteristics can be static attribute features used to describe user identity information, specifically including the target user's age, gender, and speech clarity. User behavior features are used to describe the dynamic behavior features of the user's identity information, specifically including the amplitude of movements, the frequency of movements, and lip posture. Specifically, the target user's basic characteristics can be identified based on their voiceprint features to obtain the target user's static attribute features, and the target user's behavior features during voice interaction can be identified based on their lip movement features to obtain the target user's dynamic behavior features. The static attribute features and dynamic behavior features of the target user are then fused to obtain an attribute feature set describing the target user's identity information. For example, based on the voiceprint features, if the target user's static attribute features are determined to be young, male, and with an unclear voiceprint, and it is determined that the target user is a new user with a large consumer loan need, and the obtained lip movement features are small amplitude of movements, fast frequency of movements, stiff lip posture, and a state of speech asynchrony, then the user behavior profile corresponding to this target can be determined to be a high-risk suspected fraud user in a state of anxiety. This embodiment does not impose specific limitations on this.

[0067] This technical solution can determine the user behavior profile of the target user based on the user's voiceprint features and lip movement features, realizing multi-dimensional identification of the target user, further improving the efficiency and accuracy of identifying the target user's business needs, and ensuring the service quality for the target user.

[0068] S250. Generate a business knowledge graph corresponding to the target user based on business content keywords and user behavior profiles.

[0069] This involves determining the correlation between business content keywords and the behavioral attribute features of target users in user behavior profiles, thereby generating a business knowledge graph corresponding to the target user. Specifically, based on business content keywords, the business needs information of target users can be determined. Then, based on the business needs information of target users and user behavior profiles, the correlation between the business needs information of target users and various behavioral attribute features can be determined, and a business knowledge graph corresponding to the target user can be generated accordingly.

[0070] For example, based on keywords related to business content, the target user's business needs are identified as large-amount unsecured consumer loans, fast approval, and high-amount disbursement. The target user's behavioral profile identifies them as a high-risk, suspected fraudster in a state of anxiety. Their behavioral characteristics include being young, male, having an unclear voiceprint, speaking rapidly with stiff lip movements, and frequently making business inquiries during non-working hours at night. This allows for the generation of a business knowledge graph for the target user: the target user is interested in large-amount unsecured consumer loans and urgently needs fast approval and large disbursement; the target user expresses abnormal emotions and has unclear intentions; and the target user makes business inquiries during non-working hours at night, and the number of inquiries is high. Based on this, risk control strategies for the target user can be triggered, such as secondary liveness detection, manual qualification verification, and investigation of related accounts. This embodiment does not impose specific limitations on these measures.

[0071] S260. Based on the business knowledge graph, generate a request processing strategy for the target user, and generate and feed back the processing results of the business requests initiated by the target user based on the request processing strategy.

[0072] The technical solution of this invention, in response to a business processing request initiated by a target user, acquires the target user's voice data and video data; based on the voice data, determines the target user's voice text and voiceprint features; and based on the video data, determines the target user's lip movement features to identify the text content of the voice text, obtaining business content keywords; based on the voiceprint features and lip movement features, determines the target user's user behavior profile; and based on the business content keywords and user behavior profile, generates a business knowledge graph corresponding to the target user. This implementation scheme can determine the target user's business content keywords based on the user's voice text; determine the target user's user behavior profile based on the user's voiceprint features and lip movement features, and generate a business knowledge graph for the target user accordingly. It achieves multimodal fusion of biometric recognition and visual semantic understanding, performs multi-dimensional recognition of the target user's voice text content and biometric features, improves the efficiency and accuracy of identifying the target user's business needs, enhances the accuracy and reliability of business processing results, and ensures the service quality for the target user.

[0073] Example 3

[0074] Figure 3 This is a schematic diagram of a business processing device provided in Embodiment 3 of the present invention. The business processing device provided in this embodiment of the present invention is applicable to situations where the efficiency and accuracy of identifying the business needs of the target user are low during the processing of business requests. This business processing device can be implemented in hardware and / or software, such as... Figure 3As shown, it specifically includes: a data acquisition module 310, a data parsing module 320, a map generation module 330, and a request processing module 340. Among them,

[0075] The data acquisition module 310 is used to acquire the user's voice data and user video data in response to a business processing request initiated by the target user.

[0076] The data parsing module 320 is used to determine the user voice text and user voiceprint features of the target user based on the user voice data, and to determine the lip movement features of the target user based on the user video data.

[0077] The knowledge graph generation module 330 is used to generate a business knowledge graph corresponding to the target user based on the user's voice text, the user's voiceprint features, and the lip movement features.

[0078] The request processing module 340 is used to generate a request processing strategy for the target user based on the business knowledge graph, and to generate and feed back the processing result of the business request initiated by the target user based on the request processing strategy.

[0079] This solution can determine the target user's voice text, voiceprint features, and lip movement features based on the target user's voice and video data. It then generates a business knowledge graph of the target user, thereby generating a request processing strategy for the target user's business requests and providing feedback on the request processing results. This improves the efficiency and accuracy of identifying the target user's business needs, enhances the accuracy and reliability of business processing results, and ensures the service quality for the target user.

[0080] Optionally, the graph generation module 330 is specifically used to recognize the text content of the user's voice text to obtain business content keywords;

[0081] Based on the user's voiceprint features and lip movement features, a user behavior profile of the target user is determined;

[0082] Based on the business content keywords and the user behavior profile, a business knowledge graph corresponding to the target user is generated.

[0083] Optionally, the map generation module 330 is further configured to determine the basic user characteristics of the target user based on the user's voiceprint characteristics, and to determine the user behavior characteristics of the target user based on the lip movement characteristics.

[0084] Based on the user's basic characteristics and user behavior characteristics, a user behavior profile of the target user is determined.

[0085] Optionally, the data parsing module 320 is specifically used to determine the target voice data and user voiceprint features of the target user based on the user voice data and a preset voice separation algorithm.

[0086] The target speech data is subjected to speech recognition to obtain the user's speech text.

[0087] Optionally, the data parsing module 320 is specifically used to generate at least one lip image video frame of the target user based on the user video data;

[0088] Feature extraction is performed on multiple key lip features associated with each of the lip image video frames, and the motion trajectory corresponding to each key lip feature is obtained.

[0089] Based on the motion trajectory corresponding to each of the key lip feature points, the lip movement features of the target user are determined.

[0090] Optionally, the device may also include:

[0091] The strategy adjustment module is used to determine whether the target user's request processing strategy meets the preset request processing strategy optimization requirements based on the user's voice text, user voiceprint features, and lip movement features, after generating a request processing strategy for the target user based on the business knowledge graph, generating a processing result for the business request initiated by the target user based on the request processing strategy and feeding it back.

[0092] If so, the request processing strategy is adjusted to obtain the adjusted request processing strategy.

[0093] The business processing apparatus provided in the embodiments of the present invention can execute the business processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0094] Example 4

[0095] Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0096] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 and a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 can also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0097] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0098] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as business processing methods.

[0099] In some embodiments, the business processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the business processing method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the business processing method by any other suitable means (e.g., by means of firmware).

[0100] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0101] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0102] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0104] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0105] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.

[0106] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0107] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A business processing method, characterized in that, include: In response to a business processing request initiated by a target user, the user's voice data and user video data are obtained; Based on the user voice data, determine the user voice text and user voiceprint features of the target user, and based on the user video data, determine the lip movement features of the target user; Based on the user's voice text, the user's voiceprint features, and the lip movement features, a business knowledge graph corresponding to the target user is generated; Based on the business knowledge graph, a request processing strategy for the target user is generated, and the processing result of the business request initiated by the target user is generated and fed back based on the request processing strategy.

2. The method according to claim 1, characterized in that, The step of generating a business knowledge graph corresponding to the target user based on the user's voice text, the user's voiceprint features, and the lip movement features includes: The text content of the user's voice text is recognized to obtain business content keywords; Based on the user's voiceprint features and lip movement features, a user behavior profile of the target user is determined; Based on the business content keywords and the user behavior profile, a business knowledge graph corresponding to the target user is generated.

3. The method according to claim 2, characterized in that, The step of determining the user behavior profile of the target user based on the user's voiceprint features and lip movement features includes: Based on the user's voiceprint features, the basic user characteristics of the target user are determined, and based on the lip movement features, the user behavior characteristics of the target user are determined. Based on the user's basic characteristics and user behavior characteristics, a user behavior profile of the target user is determined.

4. The method according to claim 1, characterized in that, The step of determining the target user's voice text and voiceprint features based on the user's voice data includes: Based on the user's voice data, and using a preset voice separation algorithm, the target user's target voice data and user voiceprint features are determined. The target speech data is subjected to speech recognition to obtain the user's speech text.

5. The method according to claim 1, characterized in that, The step of determining the lip movement features of the target user based on the user video data includes: Based on the user video data, generate at least one lip image video frame of the target user; Feature extraction is performed on multiple key lip features associated with each of the lip image video frames, and the motion trajectory corresponding to each key lip feature is obtained. Based on the motion trajectory corresponding to each of the key lip feature points, the lip movement features of the target user are determined.

6. The method according to claim 1, characterized in that, After generating a request processing strategy for the target user based on the business knowledge graph, and generating and feeding back the processing result of the business request initiated by the target user based on the request processing strategy, the method further includes: Based on the user's voice text, user voiceprint features, and lip movement features, determine whether the target user's request processing strategy meets the pre-set request processing strategy optimization requirements; If so, the request processing strategy is adjusted to obtain the adjusted request processing strategy.

7. A business processing apparatus, characterized in that, include: The data acquisition module is used to acquire the user's voice data and user video data in response to a business processing request initiated by the target user. The data parsing module is used to determine the user's voice text and user voiceprint features of the target user based on the user's voice data, and to determine the lip movement features of the target user based on the user's video data. The knowledge graph generation module is used to generate a business knowledge graph corresponding to the target user based on the user's voice text, the user's voiceprint features, and the lip movement features. The request processing module is used to generate a request processing strategy for the target user based on the business knowledge graph, and generate and feed back the processing result of the business request initiated by the target user based on the request processing strategy.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that is executed by the at least one processor to enable the at least one processor to perform the business processing method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the business processing method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the business processing method according to any one of claims 1-6.