Information exchange methods, devices, equipment and storage media

By analyzing and extracting features from user-sent videos, and using virtual digital humans to display response videos, the problem of insufficient diversity in intelligent customer service systems is solved, improving the flexibility of information interaction and user experience.

CN116312537BActive Publication Date: 2026-01-30INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310136090.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2026-01-30
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

Most existing intelligent customer service systems are based on text-based dialogue, resulting in insufficient flexibility in information exchange and a poor user experience, making them particularly unsuitable for the elderly, children, and people with distinct occupational characteristics.

Method used

By acquiring videos sent by users, parsing and processing them to obtain multi-frame sub-images and voice/text information, extracting features, and using virtual digital humans to display responsive videos, multimedia information interaction is supported.

Benefits of technology

It improves the flexibility of information interaction and user experience, allowing users to describe problems via video and voice, and virtual digital humans provide intuitive responses, enhancing the diversity and friendliness of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312537B_ABST
    Figure CN116312537B_ABST
Patent Text Reader

Abstract

This disclosure provides an information interaction method applicable to the fields of information processing technology, artificial intelligence technology, and fintech. The method includes: responding to an information interaction request sent by a target user, acquiring a video to be processed carried in the request; parsing the video to be processed to obtain multiple sub-frames and the target user's voice and text information; extracting features from the multiple sub-frames to obtain multiple sets of feature data; determining a response result associated with the information interaction request based on the multiple sets of feature data and the target user's voice and text information; converting the response result into a response video using a preset rendering method, and displaying the response video to the target user in the form of a virtual digital human. This disclosure also provides an information interaction device, apparatus, and storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of information processing technology, artificial intelligence technology, and financial technology, and more specifically to an information interaction method, apparatus, device, and storage medium. Background Technology

[0002] Information interaction refers to the process of sending and receiving information, which can be achieved through intelligent customer service systems. With the continuous development of information technology and artificial intelligence, intelligent customer service systems are becoming increasingly widespread. In realizing the inventive concept of this publication, the inventors discovered the following problems in related technologies: current intelligent customer service systems are generally based on text-based dialogue, but with the advancement of multimedia technology, people are no longer limited to obtaining response information by inputting text. Therefore, intelligent customer service systems in related technologies suffer from insufficient diversity, thereby reducing the flexibility of information interaction and the user experience. Summary of the Invention

[0003] In view of the above problems, this disclosure provides information interaction methods, apparatus, devices, storage media and program products that improve the flexibility of information interaction and user experience.

[0004] One aspect of this disclosure provides an information interaction method, comprising: in response to an information interaction request sent by a target user, acquiring a video to be processed carried in the information interaction request; parsing the video to be processed to obtain multiple sub-frames of images and the target user's voice and text information; extracting features from the multiple sub-frames of images to obtain multiple sets of feature data; determining a response result associated with the information interaction request based on the multiple sets of feature data and the target user's voice and text information; converting the response result into a response video using a preset rendering method, and displaying the response video to the target user in the form of a virtual digital human.

[0005] According to an embodiment of this disclosure, the step of parsing the video to be processed to obtain multiple sub-frames and the voice-text information of the target user includes: extracting the video to be processed frame by frame according to a preset extraction rate to obtain multiple sub-frames; extracting the voice information of the target user from the video to be processed using a voice extraction component; and converting the voice information into the voice-text information using a voice recognition component.

[0006] According to embodiments of this disclosure, the step of extracting features from multiple frames of the sub-images to obtain multiple sets of feature data includes: for each frame of the sub-images: extracting text information of the target pixel position from the sub-image based on preset pixel coordinates, wherein the text information includes metadata and personalized input data; and determining the feature data based on the metadata and the personalized input data.

[0007] According to embodiments of this disclosure, determining the response result associated with the information interaction request based on multiple sets of the feature data and the voice and text information of the target user includes: determining the target user's target question based on the multiple sets of the feature data and the voice and text information of the target user; searching for a target answer from a data warehouse based on the target question; and generating the response result based on the target question and the target answer.

[0008] According to embodiments of this disclosure, the method further includes: determining a plurality of related questions similar to the target question from the target repository based on the target question; sorting the plurality of related questions according to the similarity between the related questions and the target question to obtain a sorting result; taking the related question ranked first in the sorting result as the target related question; searching for a target related answer from the data warehouse based on the target related question; and displaying the target related question and the target related answer to the target user through the virtual digital human.

[0009] According to embodiments of this disclosure, the method further includes: configuring a jump key in the response video; and directly jumping to the operation interface associated with the information interaction request in response to a target user clicking the jump key.

[0010] According to embodiments of this disclosure, the method further includes: setting an input box at the bottom of the response video to allow the target user to input information; and displaying the information input by the target user in the form of a pop-up screen.

[0011] According to embodiments of this disclosure, the method further includes: configuring a virtual service room for the target user in response to the target user's request for virtual services; and rendering the background of the virtual service room and the virtual digital human in real time using virtual reality technology in the virtual service room.

[0012] Another aspect of this disclosure provides an information interaction device, comprising: an acquisition module, configured to acquire a video to be processed carried in an information interaction request sent by a target user; a processing module, configured to parse the video to be processed to obtain multiple sub-frames and the target user's voice and text information; an extraction module, configured to extract features from the multiple sub-frames to obtain multiple sets of feature data; a first determination module, configured to determine a response result associated with the information interaction request based on the multiple sets of feature data and the target user's voice and text information; and a conversion module, configured to convert the response result into a response video using a preset rendering method and display the response video to the target user in the form of a virtual digital human.

[0013] Another aspect of this disclosure provides an electronic device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the aforementioned information interaction method.

[0014] Another aspect of this disclosure provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the aforementioned information exchange method.

[0015] Another aspect of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the above-described information interaction method.

[0016] According to the information interaction method, apparatus, device, storage medium, and program product provided in the embodiments of this disclosure, the method responds to an information interaction request sent by a target user and acquires a video to be processed; parses and processes the video to be processed to obtain multiple sub-frame images and voice / text information; extracts features from the multiple sub-frame images to obtain feature data; determines a response result based on the feature data and voice / text information; converts the response result into a response video using a preset rendering method, and displays it to the target user in the form of a virtual digital human. Because the method can acquire a video to be processed during information interaction, perform parsing and feature extraction on the video to obtain a response result, and display the response video in the form of a virtual digital human, the information interaction method of this application is no longer limited to a text-based question-and-answer method that can only acquire text information. It at least partially overcomes the problem of insufficient diversity in related intelligent customer service systems, thereby achieving the technical effect of improving the flexibility of information interaction and the user experience. Attached Figure Description

[0017] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0018] Figure 1 This schematic diagram illustrates a system architecture of an information interaction method and apparatus according to embodiments of the present disclosure;

[0019] Figure 2 A flowchart illustrating an information interaction method according to an embodiment of the present disclosure is shown schematically;

[0020] Figure 3 A flowchart illustrating an information interaction method according to other embodiments of this disclosure is shown schematically;

[0021] Figure 4 A schematic diagram illustrating the structure of an information interaction device according to embodiments of the present disclosure is shown; and

[0022] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing an information interaction method according to an embodiment of the present disclosure. Detailed Implementation

[0023] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0026] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).

[0027] Current intelligent customer service systems are generally based on text-based dialogue, where users manually enter a question in text on their terminal device, and the server can search the text database based on the question, match it, and return the corresponding answer.

[0028] With the advancement of multimedia technology, users are no longer accustomed to text-based communication, especially the elderly, children, farmers, drivers, and chefs. Due to factors such as vision, age, interactivity, habits, and occupational characteristics, these groups prefer multimedia formats like images and short videos for information interaction. Therefore, in this information age, it is necessary to change the mindset of traditional text-based interaction and incorporate image and short video interaction, as well as guided digital human-narrated videos. This will enhance the flexibility of information interaction in terms of interactivity, user-friendliness, and interface design, thereby improving the user experience of intelligent customer service systems.

[0029] In view of this, embodiments of the present disclosure provide an information interaction method, apparatus, device, storage medium, and program product to improve the flexibility of information interaction and user experience. Specifically, the method may include, in response to an information interaction request sent by a target user, acquiring a video to be processed carried in the information interaction request; parsing the video to be processed to obtain multiple sub-frame images and the target user's voice and text information; extracting features from the multiple sub-frame images to obtain multiple sets of feature data; determining a response result associated with the information interaction request based on the multiple sets of feature data and the target user's voice and text information; converting the response result into a response video using a preset rendering method, and displaying the response video to the target user in the form of a virtual digital human.

[0030] It should be noted that the information interaction method and apparatus determined in the embodiments of this disclosure can be used in the fields of information processing technology, artificial intelligence technology, and financial technology, and can also be used in any field other than the fields of information processing technology, artificial intelligence technology, and financial technology. The embodiments of this disclosure do not limit the application field of the determined information interaction method and apparatus.

[0031] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of the data involved (including but not limited to user personal information, recorded videos, extracted metadata, extracted personalized data, etc.) comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.

[0032] Figure 1 A schematic diagram illustrating the system architecture of an information interaction method and apparatus according to embodiments of the present disclosure is provided.

[0033] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0034] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send information exchange requests, or send pending videos, texts, images, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as financial applications, shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0035] Terminal devices 101, 102, and 103 can be various electronic devices that support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0036] Server 105 can be a server providing various services, such as a backend management server that processes information interaction requests sent by users using terminal devices 101, 102, and 103 (this is just an example). The backend management server can analyze and process received information interaction requests, videos, and other data, and feed back the processing results (such as web pages, information, or data obtained or generated based on user information interaction) to the terminal devices. For example, server 105 can respond to an information interaction request sent by a target user by obtaining the video to be processed carried in the information interaction request; parsing the video to be processed to obtain multiple sub-frames of images and the target user's voice and text information; extracting features from the multiple sub-frames of images to obtain multiple sets of feature data; determining the response result associated with the information interaction request based on the multiple sets of feature data and the target user's voice and text information; converting the response result into a response video using a preset rendering method; and displaying the response video to the target user in the form of a virtual digital human.

[0037] It should be noted that the information interaction method provided in this embodiment can generally be executed by server 105. Correspondingly, the information interaction device provided in this embodiment can generally be located in server 105. The information interaction method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the information interaction device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Alternatively, the information interaction method provided in this embodiment can also be executed by terminal devices 101, 102, or 103, or by other terminal devices different from terminal devices 101, 102, or 103. Correspondingly, the information interaction device provided in this embodiment can also be located in terminal devices 101, 102, or 103, or in other terminal devices different from terminal devices 101, 102, or 103.

[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0039] The following will be based on Figure 1 The described system architecture, through Figures 2-3 The information interaction method of the disclosed embodiments will be described in detail.

[0040] Figure 2 A flowchart illustrating an information interaction method according to an embodiment of the present disclosure is shown schematically.

[0041] like Figure 2 As shown, the information interaction method in this embodiment includes operations S201 to S205.

[0042] In operation S201, in response to the information interaction request sent by the target user, the video to be processed carried in the information interaction request is obtained.

[0043] In operation S202, the video to be processed is parsed to obtain multiple sub-frame images and the target user's voice and text information.

[0044] In operation S203, feature extraction is performed on multiple sub-frames of images to obtain multiple sets of feature data.

[0045] In operation S204, based on multiple sets of feature data and the target user's voice and text information, the response result associated with the information interaction request is determined.

[0046] When operating S205, the response result is converted into a response video using a preset rendering method, and the response video is displayed to the target user in the form of a virtual digital human.

[0047] According to embodiments of this disclosure, the target user can be understood as a user using the aforementioned terminal device. An information interaction request can be understood as being generated when the target user encounters a problem during the process of handling business and needs to find a corresponding answer. For example, if the target user encounters a problem where querying business A fails during the process of handling business, the information interaction request could be a request to find a method to solve the problem.

[0048] According to embodiments of this disclosure, the information interaction request may carry images, text, or videos describing a problem sent by the target user. For example, when a user encounters a problem where querying service A fails, they can record the entire query process for service A and provide voice accompaniment to obtain a video to be processed. It is understood that the user can record the video themselves, or the server can record the video with the user's permission. For example, before recording the video, a request to obtain the recording video can be sent to the user, and recording will only proceed if the user agrees or authorizes recording. All of the above processes comply with relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0049] According to embodiments of this disclosure, the videos or images uploaded by the target user are compressed using a preset compression standard, such as the H.264 compression standard. This ensures that the size of the videos and images is controlled within a preset memory threshold while maintaining clarity. For example, short videos are limited to 20MB, and images to 3MB, thus saving space, improving resource utilization, and ensuring smooth data transmission. The preset compression standard and preset memory threshold can be adaptively adjusted according to actual needs.

[0050] According to embodiments of this disclosure, target users can choose to raise questions through voice dialogue, such as "I encountered a problem when logging in that the card number I entered was incorrect," or they can directly upload previously recorded videos and send them to the virtual digital human customer service.

[0051] According to embodiments of this disclosure, users can use video recording to describe problems, making the description clearer and more intuitive. Furthermore, if the video requires intervention from maintenance personnel, it facilitates their identification of the cause of the problem and the provision of corresponding solutions, thus improving business processing efficiency. In addition, users are no longer limited to text input when describing problems, reducing the complexity of manual text entry. Users can choose to input voice, send images, or send videos to describe the problems they encounter, enhancing the user experience.

[0052] According to embodiments of this disclosure, the parsing process of the video to be processed can yield multiple sub-frames and the target user's speech and text. The parsing process can be implemented using some extraction components.

[0053] According to embodiments of this disclosure, feature extraction from multiple sub-frames can be used to extract information carried in the sub-frames, such as the amount, duration, number, and error number of service A. It can also extract personalized input from the target user, such as the input service duration, number, and amount. This allows for a more accurate identification of solutions to the problem based on this feature data.

[0054] According to embodiments of this disclosure, the problem encountered by the target user can be determined by comparing and calibrating the feature data obtained from feature extraction with the target user's voice and text information, or by content splicing. For example, by extracting keywords such as "query," "A," and "failure" from the voice and text, and combining them with information such as the error number extracted from the sub-image, after verification and calibration, it can be concluded that the target user's problem is "failed to query service A." As another example, by extracting keywords such as "login" and "failure" from the voice and text, and combining them with information such as the error number and service A number extracted from the sub-image, it can be concluded that the target user's problem is "failed to log in to service A." It should be noted that the user's consent or authorization can be obtained before acquiring the user's login information. For example, a request to obtain the user's login information can be sent before login. Subsequent operations are performed only if the user agrees or authorizes the acquisition of the user's login information. All of the above processes comply with relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0055] According to embodiments of this disclosure, after identifying the problems that a target user may encounter, answers corresponding to these problems can be retrieved from a data warehouse, and the response result can consist of the analyzed problems and the answers to the responses.

[0056] According to embodiments of this disclosure, the preset rendering method may include RTC (Real-time Communication) cloud-distributed real-time generation and rendering technology. The questions and answers in the response results are rendered into a response video and displayed to the target user. The response video may be a real-time video with subtitles at the bottom, featuring a virtual digital human providing on-screen narration and accompanied by voice. A virtual digital human can be understood as a digitally created character that closely resembles a human figure.

[0057] According to embodiments of this disclosure, the response video can be compressed using a preset compression standard, such as the H.264 compression standard. Furthermore, the backend server can dynamically store the most recent preset time period of dynamic video using a caching method, allowing the target user to view or review it. The preset compression standard, caching method, and most recent preset time period can be adaptively adjusted according to actual needs.

[0058] According to the information interaction method, apparatus, device, storage medium, and program product provided in the embodiments of this disclosure, the method responds to an information interaction request sent by a target user and acquires a video to be processed; parses and processes the video to be processed to obtain multiple sub-frame images and voice / text information; extracts features from the multiple sub-frame images to obtain feature data; determines a response result based on the feature data and voice / text information; converts the response result into a response video using a preset rendering method, and displays it to the target user in the form of a virtual digital human. Because the method can acquire a video to be processed during information interaction, perform parsing and feature extraction on the video to obtain a response result, and display the response video in the form of a virtual digital human, the information interaction method of this application is no longer limited to a text-based question-and-answer method that can only acquire text information. It at least partially overcomes the problem of insufficient diversity in related intelligent customer service systems, thereby achieving the technical effect of improving the flexibility of information interaction and the user experience.

[0059] According to embodiments of this disclosure, before a target user sends an information interaction request, the target user can first enter a virtual service room. For example, in response to the target user's request for a virtual service, a virtual service room is configured for the target user; in the virtual service room, the background and virtual digital human are rendered in real time using virtual reality technology.

[0060] According to embodiments of this disclosure, when a target user requires a virtual service, a virtual service room is configured for the target user after obtaining permission to enter the virtual service. For example, if the target user opens the virtual service portal and this is their first time entering, the server can generate a virtual service room for the target user in real time and render the background and virtual digital human of the virtual service room in real time using Virtual Reality (VR) technology. The background and appearance of the virtual digital human can be set by default, and the target user can choose the background and appearance of the virtual digital human according to their interests. When the target user enters the virtual service room subsequently, the background and virtual digital human can be rendered according to the target user's interests.

[0061] According to embodiments of this disclosure, the entry point for a virtual service can be embedded in the user interface as a floating small icon, without affecting other information on the target user's browsing interface.

[0062] According to embodiments of this disclosure, the background of the virtual service room supports 3D (3D) settings such as scenery, landmarks, cartoons, and animations.

[0063] According to embodiments of this disclosure, the virtual digital human can support various image types, such as 3D images like cartoon characters and anime characters. During information interaction, the virtual digital human can semantically express facial emotions like "joy, anger, sorrow, and happiness," as well as physical actions such as shaking its head, reaching out, and spinning. The target user can change the virtual digital human's hairstyle, clothing, shoes, and accessories (such as backpacks and hair clips). After the virtual digital human has been set up according to the target user's needs, it can announce in voice, "I am your service assistant: xx, how can I help you?", facilitating the target user to send pending videos, images, or text.

[0064] According to embodiments of this disclosure, the backend server can store information such as the background and virtual digital human image of the target user in the virtual service room when it is first set up, and display them by default when the target user enters the virtual service room again.

[0065] According to embodiments of this disclosure, services are provided face-to-face by offering a dedicated virtual service room for the target user. The target user can customize the background and virtual avatar of the virtual service room based on their interests, thereby enhancing user engagement and experience.

[0066] According to an embodiment of this disclosure, operation S202 may further include the following operations: extracting the video to be processed frame by frame according to a preset extraction rate to obtain multiple sub-frame images; extracting the voice information of the target user from the video to be processed using a voice extraction component, and converting the voice information into voice-text information using a voice recognition component.

[0067] According to embodiments of this disclosure, the preset extraction rate can be adaptively adjusted according to actual needs. For example, for a video to be processed, images can be extracted at 30 frames per second to obtain multiple sub-images, which can also constitute an image stream.

[0068] According to embodiments of this disclosure, the speech extraction component may include the STRAIGHT algorithm, the Merlin speech synthesis system, etc., which can extract speech from video. The speech recognition component may be implemented using methods such as NLP (Natural Language Processing). These components can convert the extracted speech into speech-to-text, so that the speech-to-text can be used to determine the problem encountered by the target user. The speech extraction component and the speech recognition component can be adaptively adjusted according to actual needs.

[0069] According to an embodiment of this disclosure, operation S203 may further include the following operations: for each sub-image in a multi-frame sub-image: extracting text information of the target pixel position from the sub-image based on preset pixel coordinates, wherein the text information includes metadata and personalized input data; determining feature data based on the metadata and personalized input data.

[0070] According to embodiments of this disclosure, when extracting feature information from a sub-image, it can be extracted based on preset pixel coordinates, such as [x top left, Y top left, x top right, Y top right, x bottom left, Y bottom left, x bottom right, Y bottom right, and text information corresponding to the coordinates].

[0071] According to embodiments of this disclosure, the metadata and personalized information of a physical product can be identified through the extraction process of multiple sub-frames. Both metadata and personalized information can be represented in the form of [x top left, Y top left, x top right, Y top right, x bottom left, Y bottom left, x bottom right, Y bottom right, and text information corresponding to the coordinates]. The metadata portion can be understood as the part uniformly set and displayed on the system front end. For example, in the case of a loan for a physical product, the metadata portion could be a text box prompt for "loan interest," a "submit" button, etc. The personalized information portion can include system pop-ups and prompts, and the front-end input portion for the target user. For example, the personalized information section can extract the text box for "Loan Term" to obtain the input information as [50, 30, 50, 50, 60, 30, 60, 50, "30"]. The metadata section can extract the error message to obtain [20, 20, 20, 40, 60, 20, 60, 40, "Loan term cannot exceed 20 years"]. In this case, it can be assumed that the target user's input term "30" is greater than "20". Therefore, with the user's permission, the reason why this type of loan term cannot exceed 20 years can be found for the user.

[0072] According to embodiments of this disclosure, metadata may also include amount, number, name, term, error number, error information, etc., such as loan amount, loan business number, loan business name, loan term, etc. The personalized information portion can be understood as personalized data input by the target user. The target user can input the business details they wish to inquire about based on their actual situation, such as a loan term of 30 years, 25 years, or 50 years, to improve the diversity of the intelligent customer service system and enhance the user experience. The extracted metadata and feature data can be used as feature data for subsequent operations. It is understood that user consent or authorization can be obtained before extracting metadata or personalized data. For example, a request to obtain metadata or personalized data can be sent to the user before extraction. Subsequent operations are performed only after the user agrees or authorizes the acquisition of information. The above processes all comply with relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0073] According to an embodiment of this disclosure, operation S204 may further include the following operations: determining the target user's target question based on multiple sets of feature data and the target user's voice and text information; searching for the target answer from a data warehouse based on the target question; and generating a response result based on the target question and the target answer.

[0074] According to embodiments of this disclosure, the target problem can refer to a question encountered or sought by a target user, and the target answer can be an answer that can solve the target problem. Specifically, performing the aforementioned feature extraction operation on an image stream or an image sent by the target user can obtain feature data. Combining the feature data with the target user's voice / text can determine the problem encountered by the target user. For example, by verifying or calibrating the feature data obtained from the feature extraction operation with the target user's voice / text information, or by content splicing, the target problem encountered by the target user can be determined.

[0075] According to embodiments of this disclosure, a data warehouse can store various questions and answers related to various businesses or products, each business, product, question, and answer being configured with an associated identifier. The process of finding a target answer from the data warehouse can include the following operations: after determining the target question, determining the identifier of the target question based on keywords in the target question, and searching for the target answer based on the identifier. A response result can be generated by concatenating the target question and the target answer.

[0076] According to embodiments of this disclosure, the above method may further include the following operations: based on the target question, determining multiple related questions similar to the target question from a target repository; sorting the multiple related questions according to the similarity between the related questions and the target question to obtain a sorting result; taking the related question ranked first in the sorting result as the target related question; searching for the target related answer from a data warehouse according to the target related question; and displaying the target related question and the target related answer to the target user in the form of a virtual digital human.

[0077] According to embodiments of this disclosure, to improve user experience, after identifying the target problem encountered by the target user, similar questions and answers can be recommended using neural networks such as information recommendation models. The related questions are similar to the target question, and the similarity between the related questions and the target question can be determined based on keywords, business scenarios, corresponding products, frequently encountered problems, etc. The related answers are the answers corresponding to the related questions.

[0078] According to embodiments of this disclosure, after finding related questions, a set of related questions and a set of related answers can be generated based on multiple related questions and answers. By calculating the similarity between the target question and each related question, the set of related questions can be sorted, for example, in descending order of similarity. Then, the related question ranking first in the sorting is selected as the target related question to be recommended to the target user. The first number can be adjusted adaptively according to actual circumstances. For example, the top 3 or top 5 related questions can be recommended to the target user as target related questions.

[0079] According to embodiments of this disclosure, the process of calculating the similarity between the target question and each associated question may include the following operations: a language processing model can be used to assign word vectors or sentence vectors to the target question or associated questions; the similarity between the target question and associated questions is determined by calculating the distance between the word vectors or sentence vectors between the target question and the associated questions. Alternatively, BERT (Bidirectional Encoder Representation from Transformers, an unsupervised pre-trained language model for natural language processing tasks), ESIM (Enhanced Sequential Inference Model), or similar models can be used for processing.

[0080] According to embodiments of this disclosure, while recommending target-related questions to target users, the corresponding target-related answers can also be recommended to the users simultaneously. For example, based on feature data obtained from feature extraction or information about physical products, product codes can be extracted. Based on the product codes, information databases can be retrieved, and personalized data can be combined to recommend "questions you might want to know." For instance, if the metadata information extracted from video frame images is credit product A with a loan interest rate of B% and a maximum loan term of C years, and the personalized information extracted from the video frame images is the information "D years" entered in the "Loan Term" text box, then the recommended "questions you might want to know" could be: "Maximum Loan Term," "Introduction to Product A," "How is Loan Interest Calculated," etc.

[0081] According to embodiments of this disclosure, the found target question, target answer, target related question, and target related answer can all be rendered in real time using a preset rendering method to generate multiple response videos with captions at the bottom, virtual digital human narration, and voice narration. The generated response videos can be compressed using a preset compression standard, and the backend server can also store dynamic videos from the most recent preset time period.

[0082] According to embodiments of this disclosure, when displaying multiple response videos to a target user, multiple real-time videos generated by a server can be transmitted to a terminal device, and the terminal device can display text for multiple questions and matching videos.

[0083] According to embodiments of this disclosure, when displaying a response video, only the cover image may be shown to save bandwidth. The response video can be adaptively displayed at the bottom, left, right, or top of the virtual service area to facilitate selection by the target user.

[0084] According to embodiments of this disclosure, in response to a target user clicking to play a response video, the front-end page of the virtual server sends a request to the back-end to retrieve and play the digital human video. The played response video may not be displayed full-screen; the virtual digital human remains smiling within the virtual server outside the video.

[0085] According to an embodiment of this disclosure, an input box is provided at the bottom of the response video so that the target user can input information; the information input by the target user is displayed in the form of a pop-up screen.

[0086] According to embodiments of this disclosure, an input box can be configured at the bottom of the response video, where the target user can enter comments. The comments entered by the target user can then be displayed to other target users as pop-ups in subsequent videos similar to digital humans.

[0087] According to embodiments of this disclosure, introducing pop-up comments into the response video can enhance user engagement and improve user experience.

[0088] According to embodiments of this disclosure, a jump button can also be configured in the response video, and in response to the target user clicking the jump button, the user can directly jump to the operation interface associated with the information interaction request.

[0089] According to embodiments of this disclosure, a "target product menu" can be embedded at the bottom of the video, and a jump button can be carried on the right side of the target product menu. After the target user clicks the jump button, the target user can directly jump to the webpage or operation interface of the target physical product information during or after the video playback, avoiding the inconvenience of the user having to rely on memory to enter the target physical product page after exiting the current interface.

[0090] According to embodiments of this disclosure, the link corresponding to the jump button can be set based on the physical product corresponding to the response video. Alternatively, it can be determined based on user behavior data. For example, a target user might click on certain items while trying to solve a problem or conduct business. These clicks and the corresponding items can serve as behavioral data. Based on this behavioral data, the link corresponding to the jump button is generated, thereby enabling recommendations based on the user's interests.

[0091] According to embodiments of this disclosure, the links corresponding to the jump keys can access the product database and obtain relevant information through the primary and secondary transaction sections of the physical product. The granularity of the primary transaction section is greater than that of the secondary transaction section; that is, the links corresponding to the jump keys in this disclosure can directly access relevant information in the more finely granular transaction section, improving business processing efficiency and enhancing the user experience.

[0092] According to embodiments of this disclosure, if a target user wants to return to the previous operation page to modify information due to an operational error, the backend server can automatically obtain cookie (data stored on the user's local terminal) or session (object storage of attributes and configuration information required for a specific user session) information and automatically refill the personalized information previously pre-filled by the user. After updating the erroneous information, the user can perform subsequent related operations again without needing to re-enter the information, thus improving business efficiency, user experience, and the convenience of business processing. It is understood that the backend server automatically obtains information from cookies or sessions only after authorization from the target user. For example, before performing the operation, a request to obtain cookie or session information can be sent to the user. The operation is performed only if the user agrees or authorizes the acquisition. All of the above processes comply with relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0093] Figure 3 A flowchart illustrating an information interaction method according to other embodiments of this disclosure is shown.

[0094] like Figure 3 The information interaction method in this embodiment includes operations S301 to S308.

[0095] When operating S301, the target user opens the virtual service portal.

[0096] In operation S302, the target user uploads videos and pictures.

[0097] Operating S303, distributed video processing.

[0098] Operating S304, distributed knowledge graph retrieval.

[0099] The S305 is used to generate real-time video of distributed digital humans.

[0100] When operating S306, display candidate videos for "Guess You May Ask".

[0101] When operating S307, the target user clicks to play the digital human video.

[0102] When operating S308, the target user clicks to be redirected to the target page.

[0103] According to embodiments of this disclosure, the process is executed by a server deployed in the cloud, employing distributed computing to increase overall processing speed. The content between operations S301 and S308 can be referenced to the relevant content between operations S201 and S205, and will not be repeated here.

[0104] According to the embodiments provided in this disclosure, technologies such as video, digital humans, and automatic redirection to target pages can be introduced to enhance the interactivity of the question-and-answer system. The embodiments of this disclosure mainly include processes such as target user uploading videos, a distributed video processing system, real-time stitching and rendering of distributed digital human videos, a distributed knowledge graph retrieval system, displaying candidate short videos for "you might want to ask," customer selection of target videos, playback of digital human short videos, and redirection to target pages.

[0105] Distributed video processing systems can analyze video footage and audio uploaded by target users frame by frame, splice together the target user's questions, and extract physical and textual information.

[0106] The distributed knowledge graph retrieval system can match a knowledge graph database with the target user's questions and physical product information, and use artificial intelligence deep learning to match and recommend multiple related questions and return the corresponding answers.

[0107] This distributed digital human video real-time rendering system can stitch video information together in real time based on multiple related questions and answers using RTC real-time communication technology, and add subtitles at the bottom. It can also automatically match the digital human's image and voice according to the target user's characteristics. Furthermore, at the bottom of the video, the target user can click to jump to the target product's page, reducing the complexity of the target user accessing the product page from memory.

[0108] The information interaction method provided in this disclosure can offer a dedicated virtual service room for target users, providing face-to-face service support. Target users can customize the background and choose the anchor avatar of the virtual service room according to their interests. Target users can upload multimedia information such as videos and pictures, reducing the complexity of manually entering text and making the user experience more convenient. A menu-driven guide eliminates the complexity of manually navigating through service pages, especially when there are many service sections, greatly saving users' time. Introducing pop-up comments into digital human videos can enhance user experience and engagement.

[0109] It should be noted that, unless it is explicitly stated that there is a sequential order of execution between different operations, or that there is a sequential order of execution between different operations in terms of technical implementation, the execution order between multiple operations may not be significant, and multiple operations may be executed simultaneously.

[0110] Based on the above information interaction method, this disclosure also provides an information interaction device. The following will be combined with... Figure 4 The device is described in detail.

[0111] Figure 4 A schematic block diagram of an information interaction device according to an embodiment of the present disclosure is shown.

[0112] like Figure 4 As shown, the information interaction device 400 in this embodiment includes an acquisition module 410, a processing module 420, an extraction module 430, a first determination module 440, and a conversion module 450.

[0113] The acquisition module 410 is used to acquire the video to be processed carried in the information interaction request sent by the target user in response to the information interaction request.

[0114] The processing module 420 is used to parse and process the video to be processed to obtain multiple sub-frame images and the voice and text information of the target user.

[0115] The extraction module 430 is used to extract features from multiple frames of the sub-images to obtain multiple sets of feature data.

[0116] The first determining module 440 is used to determine the response result associated with the information interaction request based on multiple sets of the feature data and the voice and text information of the target user.

[0117] The conversion module 450 is used to convert the response result into a response video using a preset rendering method, and to display the response video to the target user in the form of a virtual digital human.

[0118] According to the information interaction method, apparatus, device, storage medium, and program product provided in the embodiments of this disclosure, the method responds to an information interaction request sent by a target user and acquires a video to be processed; parses and processes the video to be processed to obtain multiple sub-frame images and voice / text information; extracts features from the multiple sub-frame images to obtain feature data; determines a response result based on the feature data and voice / text information; converts the response result into a response video using a preset rendering method, and displays it to the target user in the form of a virtual digital human. Because the method can acquire a video to be processed during information interaction, perform parsing and feature extraction on the video to obtain a response result, and display the response video in the form of a virtual digital human, the information interaction method of this application is no longer limited to a text-based question-and-answer method that can only acquire text information. It at least partially overcomes the problem of insufficient diversity in related intelligent customer service systems, thereby achieving the technical effect of improving the flexibility of information interaction and the user experience.

[0119] According to embodiments of this disclosure, the processing module may further include a first extraction unit and a second extraction unit.

[0120] The first extraction unit is used to extract the video to be processed frame by frame according to a preset extraction rate to obtain multiple frames of the sub-images.

[0121] The second extraction unit is used to extract the voice information of the target user from the video to be processed using a voice extraction component, and to convert the voice information into the voice-text information using a voice recognition component.

[0122] According to embodiments of this disclosure, the extraction module may further include a third extraction unit and a first determination unit.

[0123] The third extraction unit is used to extract text information of the target pixel position from the sub-image based on preset pixel coordinates, wherein the text information includes metadata and personalized input data.

[0124] The first determining unit is used to determine the feature data based on the metadata and the personalized input data.

[0125] According to embodiments of this disclosure, the first determining module may further include a second determining unit, a searching unit, and a first generating unit.

[0126] The second determining unit is used to determine the target question of the target user based on multiple sets of feature data and the voice and text information of the target user.

[0127] The search unit is used to search for the target answer from the data warehouse based on the target question.

[0128] The first generation unit is used to generate the response result based on the target question and the target answer.

[0129] According to embodiments of this disclosure, the information interaction device may further include a second determining module, a sorting module, a third determining module, a searching module, and a first display module.

[0130] The second determining module is used to determine, based on the target problem, multiple related problems similar to the target problem from the target repository.

[0131] The sorting module is used to sort multiple related questions according to the similarity between the related questions and the target question, and obtain a sorting result.

[0132] The third determining module is used to select the association problem that ranks first in the sorting results as the target association problem.

[0133] The search module is used to search for the target-related answer from the data warehouse based on the target-related question.

[0134] The first display module is used to display the target-related question and the target-related answer to the target user through the virtual digital human.

[0135] According to embodiments of this disclosure, the information interaction device may further include a first configuration module and a jump module.

[0136] The first configuration module is used to configure jump keys in the response video.

[0137] The jump module is used to directly jump to the operation interface associated with the information interaction request in response to the target user clicking the jump button.

[0138] According to embodiments of this disclosure, the information interaction device may further include a setting module and a second display module.

[0139] The settings module is used to set an input box at the bottom of the response video so that the target user can input information.

[0140] The second display module is used to display the information input by the target user in the form of a pop-up screen.

[0141] According to embodiments of this disclosure, the information interaction device may further include a second configuration module and a rendering module.

[0142] The second configuration module is used to configure a virtual service room for the target user in response to the target user's request for a virtual service.

[0143] The rendering module is used to render the background of the virtual service room and the virtual digital human in real time using virtual reality technology.

[0144] According to embodiments of this disclosure, any multiple modules among the acquisition module 410, processing module 420, extraction module 430, first determination module 440, and conversion module 450 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the acquisition module 410, processing module 420, extraction module 430, first determination module 440, and conversion module 450 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of software, hardware, and firmware methods, or in a suitable combination of any of these methods. Alternatively, at least one of the acquisition module 410, processing module 420, extraction module 430, first determination module 440, and conversion module 450 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0145] It should be noted that the information interaction device part in the embodiments of this disclosure corresponds to the information interaction method part in the embodiments of this disclosure. The description of the information interaction device part is specifically referred to in the information interaction method part, and will not be repeated here.

[0146] Figure 5 A block diagram schematically illustrates an electronic device suitable for implementing an information interaction method according to an embodiment of the present disclosure.

[0147] like Figure 5As shown, an electronic device 500 according to an embodiment of the present disclosure includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0148] RAM 503 stores various programs and data required for the operation of electronic device 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0149] According to embodiments of this disclosure, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to a bus 504. The electronic device 500 may also include one or more of the following components connected to the I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.

[0150] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0151] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 502 and / or RAM 503 and / or one or more memories other than ROM 502 and RAM 503 described above.

[0152] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the information interaction methods provided in the embodiments of this disclosure.

[0153] When the computer program is executed by the processor 501, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0154] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 509, and / or installed from a removable medium 511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0155] In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0156] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0157] It should be noted that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of data (including but not limited to user personal information, recorded videos, extracted metadata, extracted personalized data, etc.) comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.

[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0159] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0160] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. An information interaction method, comprising: in response to an information interaction request sent by a target user, obtaining a to-be-processed video carried in the information interaction request; performing analysis processing on the to-be-processed video to obtain a plurality of sub-images and voice text information of the target user; respectively performing feature extraction on the plurality of sub-images to obtain a plurality of sets of feature data; based on the plurality of sets of feature data and the voice text information of the target user, performing proofreading, calibration or content splicing to determine a target question of the target user; based on the target question, searching for a target answer from a data warehouse; and generating a response result associated with the information interaction request according to the target question and the target answer; converting the response result into a response video by using a preset rendering method, and displaying the response video to the target user in the form of a virtual digital person, configuring a jump key in the response video, wherein a link corresponding to the jump key is set based on a corresponding real product in the response video, or is determined according to behavior data of the target user; in response to an operation of the target user clicking the jump key, directly jumping to an operation interface associated with the information interaction request, wherein the respective feature extraction on the plurality of sub-images to obtain a plurality of sets of feature data comprises: for each of the plurality of sub-images: based on a preset pixel coordinate, extracting text information to a target pixel position from the sub-image, wherein the text information includes metadata and personalized input data; based on the metadata and the personalized input data, determining the feature data. The analysis processing on the to-be-processed video to obtain a plurality of sub-images and voice text information of the target user comprises:

2. The method of claim 1, wherein, frame extraction of the to-be-processed video at a preset extraction rate to obtain a plurality of sub-images; extracting voice information of the target user from the to-be-processed video by using a voice extraction component, and converting the voice information into the voice text information by using a voice recognition component.

3. The method of claim 1, further comprising: based on the target question, determining a plurality of associated questions similar to the target question from the data warehouse; sorting a plurality of the associated questions according to a similarity between the associated questions and the target question to obtain a sorting result; taking a first number of associated questions in the sorting result as target associated questions; based on the target associated questions, searching for a target associated answer from the data warehouse; displaying the target associated questions and the target associated answer to the target user in the form of a virtual digital person.

4. The method of claim 1, further comprising: setting an input box at the bottom of the response video to enable the target user to input information; displaying the information input by the target user in the form of a pop-up screen.

5. The method of claim 1, further comprising: in response to an operation of the target user needing virtual service, configuring a virtual service room for the target user. ​ In the virtual service room, a background of the virtual service room and the virtual digital person are rendered in real time through virtual reality technology.

6. An information interaction device, comprising: an acquisition module configured to acquire a to-be-processed video carried in an information interaction request sent by a target user in response to the information interaction request; a processing module configured to perform analysis processing on the to-be-processed video to obtain a plurality of sub-images and voice text information of the target user; an extraction module configured to perform feature extraction on the plurality of sub-images respectively to obtain a plurality of sets of feature data; a first determination module configured to determine a target question of the target user based on the plurality of sets of feature data and the voice text information of the target user, find a target answer from a data warehouse based on the target question, and generate a response result associated with the information interaction request according to the target question and the target answer; a conversion module configured to convert the response result into a response video by using a preset rendering method, and display the response video to the target user in the form of a virtual digital person; a first configuration module configured to configure a jump key in the response video; and a jump module configured to jump directly to an operation interface associated with the information interaction request in response to an operation of the target user clicking the jump key. The extraction module further includes a third extraction unit and a first determination unit. The third extraction unit is configured to extract text information to a target pixel position from the sub-image based on a preset pixel coordinate, wherein the text information includes metadata and personalized input data. The first determination unit is configured to determine the feature data based on the metadata and the personalized input data.

7. An electronic device, comprising: one or more processors; a storage device configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1-5.

8. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1-5.

9. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5. ​

Citation Information

Patent Citations

  • Information interaction method and device, electronic equipment, medium and program product

    CN113392201A

  • Fault processing method and device, equipment and storage medium

    CN115423226A