Traffic police digital human interaction method and device
By designing a traffic police digital human interaction method, using multi-modal input and large language model to generate replies, combining speech synthesis and speaking face generation technology, the problem of lack of a digital human driving test training system in the existing technology is solved, and an efficient and personalized teaching experience is achieved.
Patent Information
- Application Number
- CN202510074718.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology lacks a driving test training system with digital human-based carriers, and cannot provide an interactive and personalized teaching experience.
A traffic police digital human interaction method is designed. By collecting traffic-related knowledge, building a traffic knowledge base, and using multi-modal input information (document, voice, picture) to extract problem content, calling large language models to generate reply information, combining speech synthesis and speaking face generation technology based on neural radiation field, traffic police image video is rendered to realize real-time streaming display.
It realizes low-latency and high-efficiency interaction of multi-modal content input, provides comprehensive knowledge coverage in the transportation field, supports speech synthesis in a variety of Chinese dialects, and generates realistic digital human videos, which significantly improves the teaching experience and reply quality.
Smart Images

Figure CN119938854A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of generative artificial intelligence technology, and in particular to a traffic police digital human interaction method and device. Background Art
[0002] With the continuous development of generative artificial intelligence (AI), digital human technology has gradually become a hot topic in the field of artificial intelligence. Digital human technology is a computer-generated virtual character that can simulate human appearance, language, movement, and other characteristics, and has strong interactivity and communication capabilities. Driven by new technologies such as large language models and speech synthesis, digital human interaction systems have made significant progress. By using large models to generate customized text responses, combined with highly human-like speech synthesis and lip-synced digital human videos, they can provide an excellent user experience. Digital humans are now widely used in various fields, such as customer service, education, and live streaming. However, in the field of driving test training, there is no teaching system based on digital humans.
[0003] Therefore, this application provides a traffic police digital human interaction method and device. Summary of the Invention
[0004] The technical problem to be solved by the present invention is that there is no teaching system using digital humans as carriers in the prior art. Therefore, a method and device for interacting with a digital human for traffic police are provided. The method for interacting with a digital human for traffic police comprises: Collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the organized question-answer pair format to build a traffic knowledge base; Obtaining user multimodal input information and extracting question content information from the user multimodal input information; Retrieve question-related information from the traffic knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; Convert question answer information into voice audio information; The traffic police image video is rendered according to the voice audio information through the neural radiation field-based speaking face generation technology; the rendered traffic police image video is displayed in real-time streaming.
[0005] Optionally, the obtaining of user multimodal input information and extracting question content information from the user multimodal input information includes: Obtaining user multimodal input information, and determining whether the user multimodal input information is specifically document format information, image format information, or voice format information; When the document format information is obtained, Langchain's document loading library is used to extract the problem content information in the document format information; When the image format information is obtained, use Base64 to encode the image and generate the corresponding Base64 encoded string; The voice format information in WAV format is obtained through WebRTC, the voice is converted into text format information through the SenseVoice speech recognition model, and the question content information in the text format information is extracted.
[0006] Optionally, the step of retrieving question-related information from a traffic knowledge base based on the question content information and generating question-answering information by calling a large language model based on the question-related information includes: The question text in the question-answer pair is converted into a vector representation through vectorization technology, and the vector ID corresponding to the question text is recorded. The vector is stored as a data index in the PostgresSQL vector database, and the question-answer pair text corresponding to the question text is stored in the MangoDB database. The two databases are linked through the vector ID; Obtain search content information, extract the question text from the search content information, vectorize the question text to generate a question vector, perform vector search on the question vector using the PG Vector plug-in, measure the distance between the retrieved vectors using cosine similarity, and generate a similarity score for each measurement result; sort the similarity scores, take the vector ID corresponding to the highest value, and query the question-answer text corresponding to the vector ID in the MangoDB database.
[0007] Optionally, before retrieving question-related information from a vector knowledge base based on the question content information and calling a large language model based on the question-related information to generate question reply information, the process further includes: Obtain the question content information and determine whether the question content information belongs to the transportation field. If so, proceed to the subsequent steps to retrieve question-related information from the transportation knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; if not, directly generate a guiding response to guide the topic to the transportation field.
[0008] Optionally, the method further includes: converting the question reply information into text information, and displaying the text information in real-time streaming format.
[0009] Optionally, obtaining digital human setting information, the digital human setting information including voice selection information, sentence minimum word count information, and gesture control information; Select the dialect used by the digital person and the gender of the digital person according to the voice selection information; Control the minimum length of the split sentences based on the minimum word count information of the clauses; Control the digital human's expression and movements based on the posture control information; The traffic police image video is rendered using the speaking face generation technology based on the neural radiation field in conjunction with the voice audio information.
[0010] Optionally, converting the question reply information into voice and audio information; rendering a traffic police image video using a speaking face generation technology based on a neural radiation field according to the voice and audio information; and displaying the rendered traffic police image video in real-time streaming specifically includes: Receive question reply information, and use punctuation marks as delimiters to split the text in the question reply information into sentence queues; Call the Microsoft Azure Speech Synthesis API sentence by sentence to convert text into speech audio and generate a speech queue; Using the neural radiation field-based speaking face generation technology, lip-synced digital human video frames are generated from the speech queue. Referring to the posture control parameters, the FFmpeg tool is used to synthesize the video clips and establish a video queue. Whenever a new video is added to the video sequence, the new video and the corresponding text are sent to the front end, which plays the video in real time and streams the text results. A second aspect of the present invention further provides a traffic police digital human interaction device, the traffic police digital human interaction device comprising: The database construction module is used to collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the content of the organized question-answer pair format structure to build a traffic knowledge base; A multimodal content parsing module is used to obtain user multimodal input information and extract question content information from the user multimodal input information; The knowledge base question answering module is used to retrieve question-related information from the vector knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; A speech synthesis module, used to convert question response information into speech audio information; The digital human module is used to render traffic police image videos based on speech audio information using a speaking face generation technology based on neural radiation fields; The interactive module is used to display the rendered traffic police image video in real-time streaming. A third aspect of the present invention further provides an electronic device, comprising a memory and at least one processor, wherein the memory stores instructions and data; The at least one processor calls the instructions and data in the memory so that the electronic device executes each step of the traffic police digital human interaction method as described in any one of the above items. The fourth aspect of the present invention further provides a readable storage medium, on which instructions and data are stored. When the instructions are executed by a processor, the various steps of the traffic police digital human interaction method as described above are implemented.
[0011] The implementation of the present invention has the following beneficial effects: 1. This invention supports multimodal content input, including documents, voice, and images, providing diverse user interaction methods. Answer text and digital human video are streamed to the front-end page, enabling real-time display of results without having to wait for all content to be generated. This system offers low latency and high efficiency.
[0012] 2. The traffic knowledge base constructed by the present invention includes Chinese traffic laws and regulations, test questions for subjects one and four, and test skills for subjects two and three, achieving comprehensive coverage of knowledge in the traffic field. Through retrieval enhancement generation technology, it can find the most relevant knowledge content for user questions, and provide feedback on answers after summarizing them through a large model, significantly improving the quality of responses and providing professional learning assistance for driving test candidates.
[0013] 3. The present invention supports speech synthesis of multiple Chinese dialects, and can provide voice services that are closer to users, with both intimacy and entertainment.
[0014] 4. The present invention adopts speaking face generation technology based on neural radiation field, which can generate lip-synchronized digital human videos, support posture control, and provide more realistic digital human visual effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of a traffic police digital human interaction method according to a first embodiment of the present invention; Figure 2 This is a flow chart of the traffic police digital human interaction method provided by the second embodiment of the present invention; Figure 3 This is a flow chart of a traffic police digital human interaction method provided by a third embodiment of the present invention; Figure 4 This is a structural diagram of the traffic police digital human interaction device provided by the present invention; Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] An embodiment of the present invention provides a traffic police digital human interaction method, comprising collecting a data set of traffic-related knowledge, processing the traffic-related knowledge information into a question-answer pair format, and using the content of the organized question-answer pair format structure to construct a traffic knowledge base; obtaining user multimodal input information, extracting question content information from the user multimodal input information; retrieving question-related information from the traffic knowledge base based on the question content information, calling a large language model to generate question reply information based on the question-related information; converting the question reply information into voice and audio information; rendering a traffic police image video based on the voice and audio information using a speaking face generation technology based on a neural radiation field; and displaying the rendered traffic police image video in real-time streaming mode. The present invention solves the problem in the prior art of lacking a teaching system using a digital human as a carrier.
[0017] The terms "first," "second," "third," "fourth," and so on (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar items and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that shown or described in this specification. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product, or apparatus.
[0018] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 The first embodiment of the traffic police digital human interaction method in the embodiment of the present invention includes: Collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the organized question-answer pair format to build a traffic knowledge base; Obtaining user multimodal input information and extracting question content information from the user multimodal input information; Specifically, it includes obtaining user multimodal input information and determining whether the user multimodal input information is specifically document format information, image format information or voice format information; When the document format information is obtained, Langchain's document loading library (document_loaders) is used to extract the problem content information from the document format information; When obtaining image format information, such as JPG, PNG, BMP, WEBP, etc., use Base64 to encode the image and generate the corresponding Base64 encoded string; Acquire voice information in WAV format through WebRTC (Web Real-Time Communications), convert the voice into text format using the SenseVoice speech recognition model, and extract the question content from the text format information; It can support multimodal content input including documents, voice, and pictures, and can provide diverse user interaction methods.
[0019] Retrieve question-related information from the traffic knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; Convert the question reply information into voice and audio information; another solution is also provided in this embodiment, which directly converts the question reply information into text information, and displays the text information in real-time streaming mode, which can achieve multi-mode streaming output and has the advantages of low latency and high efficiency.
[0020] The traffic police image video is rendered according to the voice audio information through the neural radiation field-based speaking face generation technology; the rendered traffic police image video is displayed in real-time streaming. See also Figure 2 The second embodiment of the traffic police digital human interaction method in the embodiment of the present invention includes: Collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the organized question-answer pair format to build a traffic knowledge base; Obtaining user multimodal input information and extracting question content information from the user multimodal input information; Specifically, it includes obtaining user multimodal input information and determining whether the user multimodal input information is specifically document format information, image format information or voice format information; When the document format information is obtained, Langchain's document loading library (document_loaders) is used to extract the problem content information from the document format information; When obtaining image format information, such as JPG, PNG, BMP, WEBP, etc., use Base64 to encode the image and generate the corresponding Base64 encoded string; Acquire voice information in WAV format through WebRTC (Web Real-Time Communications), convert the voice into text format using the SenseVoice speech recognition model, and extract the question content from the text format information; It can support multimodal content input including documents, voice, and pictures, and can provide diverse user interaction methods.
[0021] Obtain the question content information and determine whether the question content information belongs to the transportation field. If so, proceed to the subsequent steps to retrieve question-related information from the transportation knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; if not, directly generate a guided response; the guided response guides the topic direction to the transportation field.
[0022] Retrieve question-related information from the traffic knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; Convert the question reply information into voice and audio information; another solution is also provided in this embodiment, which directly converts the question reply information into text information, and displays the text information in real-time streaming mode, which can achieve multi-mode streaming output and has the advantages of low latency and high efficiency.
[0023] The traffic police image video is rendered according to the voice audio information through the neural radiation field-based speaking face generation technology; the rendered traffic police image video is displayed in real-time streaming. See also Figure 3 The third embodiment of the traffic police digital human interaction method in the embodiment of the present invention includes: Collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the organized question-answer pair format to build a traffic knowledge base; Obtaining user multimodal input information and extracting question content information from the user multimodal input information; Specifically, it includes obtaining user multimodal input information and determining whether the user multimodal input information is specifically document format information, image format information or voice format information; When the document format information is obtained, Langchain's document loading library (document_loaders) is used to extract the problem content information from the document format information; When obtaining image format information, such as JPG, PNG, BMP, WEBP, etc., use Base64 to encode the image and generate the corresponding Base64 encoded string; Acquire voice information in WAV format through WebRTC (Web Real-Time Communications), convert the voice into text format using the SenseVoice speech recognition model, and extract the question content from the text format information; It can support multimodal content input including documents, voice, and pictures, and can provide diverse user interaction methods.
[0024] Obtain the question content information and determine whether the question content information belongs to the transportation field. If so, proceed to the subsequent steps to retrieve question-related information from the transportation knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; if not, directly generate a guided response; the guided response guides the topic direction to the transportation direction.
[0025] Retrieve question-related information from the traffic knowledge base based on the question content information, and call the large language model based on the question-related information to generate question response information; The specific process of calling the large language model to generate question response information includes: The question text in the question-answer pair is converted into a vector representation through the Embedding model. The vector ID corresponding to the question text is recorded and stored as a data index in the PostgresSQL vector database. The question-answer pair text corresponding to the question text is stored in the MangoDB database. The two databases are linked through the vector ID. Obtain the search content information, extract the question text from the search content information, vectorize the question text to generate a question vector, perform vector search on the question vector using the PG Vector plug-in, measure the distance between the retrieved vectors using cosine similarity, and generate a similarity score for each measurement result; sort the similarity scores, take the vector ID corresponding to the highest value, query the question and answer text corresponding to the vector ID in the MangoDB database, and add the question and answer text to the conversation.
[0026] Convert the question reply information into voice and audio information; another solution is also provided in this embodiment, which directly converts the question reply information into text information, and displays the text information in real-time streaming mode, which can achieve multi-mode streaming output and has the advantages of low latency and high efficiency.
[0027] The traffic police image video is rendered according to the voice audio information through the neural radiation field-based speaking face generation technology; the rendered traffic police image video is displayed in real-time streaming.
[0028] Before the digital human is generated, the digital human setting information sent by the front end can be obtained, and the digital human setting information includes voice selection information, sentence minimum word number information and posture control information; Select the dialect used by the digital person and the gender of the digital person according to the voice selection information; Control the minimum length of the split sentences based on the minimum word count information of the clauses; Control the digital human's expression and movements based on the posture control information; The traffic police image video is rendered using the speaking face generation technology based on the neural radiation field in conjunction with the voice audio information.
[0029] The specific process of converting the question response information into voice and audio information; rendering the traffic police image video based on the voice and audio information using the speaking face generation technology based on the neural radiation field; and displaying the rendered traffic police image video in real-time streaming includes: Receive question reply information, and use punctuation marks as delimiters to split the text in the question reply information into sentence queues; Call the Microsoft Azure Speech Synthesis API sentence by sentence to convert text into speech audio and generate a speech queue; Using the neural radiation field-based speaking face generation technology, lip-synced digital human video frames are generated from the speech queue. Referring to the posture control parameters, the FFmpeg tool is used to synthesize the video clips and establish a video queue. Whenever a new video is added to the video sequence, the new video and the corresponding text are sent to the front-end, which plays the video in real time and streams the resulting text. This allows for real-time display of results without having to wait for all content to be generated. The system boasts low latency and high efficiency. Using neural radiation field-based speaking face generation technology, it can generate lip-synced digital human videos, support gesture control, and provide more realistic digital human visuals. The above describes the traffic police digital human interaction method in the embodiment of the present invention. The following describes the traffic police digital human interaction device in the embodiment of the present invention. Figure 4 In the embodiment of the present invention, the digital human interaction device for traffic police includes the following for the above embodiment: Database construction module 401 is used to collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the content of the organized question-answer pair format structure to build a traffic knowledge base; The multimodal content parsing module 402 is used to obtain the user's multimodal input information and extract the question content information in the user's multimodal input information; The knowledge base question answering module 403 is used to retrieve question-related information from the vector knowledge base based on the question content information, and call the large language model to generate question answer information based on the question-related information; The speech synthesis module 404 is used to convert the question answer information into speech audio information; The digital human module 405 is used to render a traffic police image video based on the speech audio information by using a speaking face generation technology based on the neural radiation field; The interactive module 406 is used to display the rendered traffic police image video in real-time streaming format. above Figure 4 The traffic police digital human interaction device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The electronic device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0030] Figure 5 The figure is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. This electronic device 500 may vary significantly depending on its configuration or performance. It may include one or more processors 510 (e.g., one or more processors), memory 520, and one or more storage media 530 (e.g., one or more storage devices, including RAM, FLASH, etc.) that store application programs 533 or data 532. The memory 520 and storage medium 530 may be either transient or persistent storage. The program stored in the storage medium 530 may include one or more modules (not shown), each of which may include a series of instructions for operating on the electronic device 500. Furthermore, the processor 510 may be configured to communicate with the storage medium 530 to execute the series of instructions stored in the storage medium 530 on the electronic device 500.
[0031] The electronic device 500 may further include one or more power supplies 540, one or more input / output interfaces 550, and / or one or more operating systems 531, such as FreeRTOS, Android, etc. It will be understood by those skilled in the art that Figure 5 The illustrated electronic device structure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0032] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions. When the instructions are run on a computer, the computer executes the steps of the traffic police digital human interaction method for logistics personnel. Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0033] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, mobile device, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0034] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A traffic police digital human interaction method, characterized in that: include: Collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the organized question-answer pair format content to build a traffic knowledge base; Obtaining user multimodal input information and extracting question content information from the user multimodal input information; Retrieve question-related information from the traffic knowledge base based on the question content information, and call the large language model to generate question response information based on the question-related information; Convert question answer information into voice audio information; Render the traffic police image video based on the speech audio information through the neural radiation field-based speaking face generation technology; The rendered traffic police image video is displayed in real-time streaming.
2. The traffic police digital human interaction method according to claim 1, characterized in that: The step of obtaining the user's multimodal input information and extracting the question content information from the user's multimodal input information includes: Obtaining user multimodal input information, and determining whether the user multimodal input information is specifically document format information, image format information, or voice format information; When the document format information is obtained, use Langchain's document loading library to extract the problem content information in the document format information; When the image format information is obtained, use Base64 to encode the image and generate the corresponding Base64 encoded string; The voice format information in WAV format is obtained through WebRTC, the voice is converted into text format information through the SenseVoice speech recognition model, and the question content information in the text format information is extracted.
3. The traffic police digital human interaction method according to claim 1, characterized in that: The retrieving question related information from the traffic knowledge base according to the question content information, and calling the large language model to generate question reply information according to the question related information includes: The question text in the question-answer pair is converted into a vector representation through vectorization technology, and the vector ID corresponding to the question text is recorded. The vector is stored in the PostgresSQL vector database as a data index, and the question-answer pair text corresponding to the question text is stored in the MangoDB database. The two databases are associated through the vector ID; Get the search content information, extract the question text in the search content information, vectorize the question text to generate a question vector, perform vector search on the question vector through the PG Vector plug-in, use cosine similarity to measure the distance of the retrieved vectors, and generate a similarity score for each measurement result; sort by similarity score, take the vector ID corresponding to the highest value, and query the question and answer text corresponding to the vector ID in the MangoDB database.
4. The traffic police digital human interaction method according to claim 3 is characterized in that: Before retrieving question related information from the vector knowledge base according to the question content information and calling the large language model to generate question reply information according to the question related information, the method further includes: Obtain the question content information and determine whether the question content information belongs to the transportation field. If so, enter the subsequent steps to retrieve question-related information from the transportation knowledge base based on the question content information, and call the large language model to generate question reply information based on the question-related information; if not, directly generate a guiding reply to guide the topic to the transportation field.
5. The traffic police digital human interaction method according to claim 1, characterized in that: Also includes: The question reply information is converted into text information, and the text information is displayed in real-time streaming.
6. The traffic police digital human interaction method according to claim 1, characterized in that: Also includes: Acquiring digital human setting information, wherein the digital human setting information includes voice selection information, sentence minimum word count information, and gesture control information; Selecting the dialect used by the digital person and the gender of the digital person according to the voice selection information; Control the minimum length of the split sentences based on the minimum word count information of the sentences; Control the digital human's expression and movements based on the posture control information.
7. The traffic police digital human interaction method according to claim 6, characterized in that: The question answer information is converted into voice and audio information; the traffic police image video is rendered according to the voice and audio information by using the speaking face generation technology based on the neural radiation field; The real-time streaming display of the rendered traffic police image video specifically includes: Receive question reply information, and divide the text in the question reply information into sentence queues using punctuation marks as delimiters; Call the Microsoft Azure speech synthesis interface sentence by sentence to convert text into speech audio and generate a speech queue; Using the neural radiation field-based talking face generation technology, lip-synchronized digital human video frames are generated from the speech queue, and the gesture control parameters are referenced to synthesize the video clips using the FFmpeg tool and establish a video queue; Whenever a new video is added to the video sequence, the new video and the corresponding text are sent to the front end, which plays the video in real time and streams the text results.
8. A traffic police digital human interaction device, characterized in that: include: The database construction module is used to collect a data set of traffic-related knowledge, process the traffic-related knowledge information into a question-answer pair format, and use the content of the organized question-answer pair format structure to build a traffic knowledge base; A multimodal content parsing module is used to obtain user multimodal input information and extract question content information from the user multimodal input information; The knowledge base question-answering module is used to retrieve question-related information from the vector knowledge base based on the question content information, and call the large language model to generate question response information based on the question-related information; A speech synthesis module, used to convert question answer information into speech audio information; Digital human module, used to render traffic police image video based on speech audio information through neural radiation field-based speaking face generation technology; The interactive module is used to display the rendered traffic police image video in real-time streaming mode.
9. An electronic device, comprising a memory and at least one processor, characterized in that: Instructions and data are stored in the memory; The at least one processor calls the instructions and data in the memory so that the electronic device executes each step of the traffic police digital human interaction method as described in any one of claims 1-7.
10. A readable storage medium having instructions and data stored thereon, characterized in that: When the instructions are executed by the processor, the various steps of the traffic police digital human interaction method as claimed in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
AI ancient poetry multi-round spoken language dialogue method and device and electronic equipment
CN121681767A