AI digital human intelligent photo frame knowledge question-answering system

By integrating voice recognition and natural language processing, the AI ​​digital human smart photo frame enables multimodal output and multi-turn dialogue, solving the problem of the single interaction mode of existing digital photo frames, improving user experience and expanding application scenarios.

CN121328701APending Publication Date: 2026-01-13ZHONGAN NET VISION (XIAMEN) YUAN UNIVERSE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511246080.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing digital photo frames have a one-way interaction method, lack the ability to proactively initiate inquiries or engage in real-time dialogue, have limited functionality, lack a built-in structured knowledge base, cannot meet users' needs for in-depth exploration of specific cultural themes, have insufficient intelligence, and have a limited display method, making them difficult to apply to occasions such as museums that require knowledge-guided tours and cultural output.

Method used

The AI-powered digital human smart photo frame integrates speech recognition, natural language processing, and knowledge base modules. It uses a digital human image to output multimodal data, enabling real-time question answering and multi-turn dialogue. It combines text, images, and voice to present content and has a built-in professional knowledge base to support in-depth interaction.

Benefits of technology

It enhances user engagement and interactive experience, transforming the product from a simple photo display device into a professional knowledge service terminal, breaking through the limitations of traditional application scenarios, and is widely used in museums, educational institutions, smart homes and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121328701A_ABST
    Figure CN121328701A_ABST
Patent Text Reader

Abstract

The invention discloses an AI digital human intelligent photo frame knowledge question-answering system which comprises an intelligent photo frame. The knowledge base management module is arranged in the storage unit and is used for storing and managing knowledge data about historical character knowledge; the digital human presentation module is arranged in the processor, is displayed on the touch display screen and is used for generating and driving a digital human image to perform content display; and the AI interaction module is arranged in the processor and is used for processing the input of the user and extracting the corresponding response content from the knowledge base management module or acquiring the corresponding response content by connecting the network communication unit with the cloud server, displaying and audio broadcasting are performed on the touch display screen by calling the digital human presentation module, and the knowledge base management module is updated at the same time. According to the AI digital human intelligent photo frame knowledge question-answering system, a hardware carrier and intelligent interaction are deeply fused, and the AI digital human intelligent photo frame knowledge question-answering system has real-time question-answering, knowledge service and dynamic display capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart device technology, and in particular to an AI digital human smart photo frame knowledge question and answer system. Background Technology

[0002] Currently, digital photo frames are quite common, with most offering features such as digital photo display, loop playback, remote synchronization, and basic cloud album management. However, existing digital photo frames still have significant limitations in their technical architecture and functional design. First, their interaction is primarily one-way content display; users can only passively view preset images or videos, lacking the ability to actively initiate inquiries or engage in real-time dialogue, thus failing to meet users' needs for in-depth exploration of specific cultural themes (such as historical figures or events). Second, their functionality is relatively limited, lacking a built-in structured professional knowledge base, especially lacking systematic and multi-dimensional knowledge data on specific historical figures like Person A, making it difficult for them to fulfill their role in cultural education and knowledge dissemination.

[0003] Furthermore, existing products suffer from significant shortcomings in terms of intelligence. Most digital photo frames fail to integrate advanced natural language processing technology and artificial intelligence interaction engines, making them unable to accurately understand complex questions posed by users in natural language or to engage in multi-turn dialogues based on context. In terms of content presentation, existing devices typically only support the playback of static images or pre-recorded videos, resulting in a limited and unengaging display method. In particular, they lack the ability to use dynamic digital human figures to explain content and express emotions, leading to a rather monotonous user experience.

[0004] Due to the aforementioned technical limitations, the application scenarios for existing digital photo frame products are severely restricted. They are difficult to use in museums, memorial halls, and other settings requiring knowledge-guided tours and cultural dissemination, and also cannot serve as effective family cultural education tools. Therefore, there is an urgent need for an AI digital human-based smart photo frame knowledge-based question-and-answer system that can deeply integrate hardware and intelligent interaction, possessing real-time question-and-answer, knowledge service, and dynamic display capabilities. Summary of the Invention

[0005] The purpose of this invention is to provide an AI digital human intelligent photo frame knowledge question and answer system.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: An AI-powered digital human-based smart photo frame knowledge-based question-and-answer system includes: A smart photo frame includes a frame, a touch display screen, a processor, a storage unit, an audio input / output unit, and a network communication unit. The frame is placed on a wall, and the touch display screen is installed inside the frame for user interaction. The processor, storage unit, audio input / output unit, and network communication unit are integrated on the back of the touch display screen and are connected to the touch display screen respectively. The knowledge base management module, which is built into the storage unit, is used to store and manage knowledge data about historical figures. The digital human presentation module, built into the processor and displayed on the touch screen, is used to generate and drive the digital human image to display content; The AI ​​interaction module, built into the processor, is used to process user input and extract corresponding response content from the knowledge base management module or obtain corresponding response content by connecting to the cloud server through the network communication unit. It then displays the information and plays audio on the touch screen by calling the digital human presentation module, while simultaneously updating the knowledge base management module.

[0007] Furthermore, the knowledge base management module contains pre-input historical figure knowledge, corresponding character lip-sync data, and corresponding historical figure images. Each historical figure knowledge corresponds to a figure introduction and at least one example. Each historical figure corresponds to a digital human image, which is composed of historical figure photos. Its lip-sync has multiple displacement points, and the lip-sync dynamically changes by moving the displacement points. When the AI ​​interaction module obtains new examples of historical figures, it adds the new examples to the database of that historical figure.

[0008] Furthermore, the audio input / output unit includes a microphone array and an audio player. The microphone array collects human voices from the surrounding environment and sends them to the processor for analysis to obtain the words in the human voices. These words are then displayed one by one on the touch screen. The AI ​​interaction module analyzes the keywords from the knowledge base management module and extracts corresponding examples of the historical figures. The corresponding examples are then played through the audio player.

[0009] Furthermore, the digital human presentation module matches and obtains the corresponding lip-sync data based on the fields to be played, arranges them according to the field order, displays the images of the corresponding historical figures on the touch screen, identifies the lip-sync positions in the images, and dynamically changes the lip-sync based on the lip-sync data.

[0010] Furthermore, the AI ​​interaction module includes a speech recognition module, a natural language understanding unit, and a question-and-answer generation unit. The speech recognition module identifies human voice data collected by the microphone array and converts the human voice data into audio text. The natural language understanding unit parses the audio text and the text input through the touch screen to obtain the user's semantics and intent, and displays the audio text and the input text on the touch screen. The question-and-answer generation unit matches corresponding historical figure knowledge and historical figure photos within the knowledge base management module based on semantics and intent.

[0011] Furthermore, the content matching between the question-and-answer generation unit and the knowledge base management module involves extracting keywords from the audio or text. First, the data within the knowledge base management module is quickly matched and extracted using these keywords. If a match fails, the system communicates with the cloud server via a network communication unit, sending the corresponding request. The cloud server then performs a semantic similarity search on the cloud knowledge base. If a match is successful, the corresponding data is sent back to the knowledge base management module. If the cloud knowledge base still fails to match, the system retrieves the vector database from the cloud server. If a satisfactory result is still not obtained, a large language model is used to generate the answer.

[0012] Furthermore, the question-and-answer generation unit includes audio data and text data. It generates questions associated with the current text data and prepares example data and lip-reading data corresponding to each question. The text of the question is displayed on the touch screen and the audio is played. Users can click on the corresponding question on the touch screen to enter the next stage of the corresponding example display and playback, or repeat the dialogue to the question. The AI ​​interaction module identifies the keywords in the dialogue, matches the question, and enters the next stage of the corresponding example display and playback.

[0013] Furthermore, the AI ​​interaction module also includes a camera, which captures the user's location and image in the current environment, identifies changes in their lip movements, determines the user speaking in the image, identifies and locks their voiceprint, and only collects the voice data of that voiceprint during the current interaction phase.

[0014] Furthermore, the AI ​​interaction module includes a voice output mode for at least one language, which is matched according to the language of the user currently speaking.

[0015] Furthermore, the network communication unit includes a wireless communication module and a wired communication module.

[0016] By adopting the above technical solution, the present invention has the following advantages compared with the prior art: 1. This invention upgrades static information display to dynamic and emotional digital human explanation by introducing digital human image-driven technology and multimodal output. It combines text, images and voice to present content, overcoming the shortcomings of existing devices that have a single display format and lack of appeal, and greatly improving the vividness and effectiveness of knowledge dissemination.

[0017] 2. By integrating speech recognition, natural language processing, and knowledge base retrieval modules, this invention enables smart photo frames to have real-time question-and-answer and multi-turn dialogue capabilities, solving the technical problem that traditional digital photo frames can only passively display content and cannot actively interact, thus significantly enhancing user participation and interactive experience.

[0018] 3. This invention transforms the product from a simple photo display device into a professional knowledge service terminal, breaking through the limitations of its traditional application scenarios. It can be widely used in fields such as museums, educational institutions, and smart homes that require cultural dissemination and knowledge explanation. Attached Figure Description

[0019] Figure 1 This is a system block diagram of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0021] It should be noted that in this invention, the terms "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. are all based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this invention and simplifying the description, and are not intended to indicate or imply that the device or element of this invention must have a specific orientation, and therefore should not be construed as a limitation of this invention. Example

[0022] refer to Figure 1 As shown, this invention discloses an AI digital human intelligent photo frame knowledge question-and-answer system, comprising: The smart photo frame includes a frame, a touch display screen, a processor, a storage unit, an audio input / output unit, and a network communication unit. The frame is placed on the wall, and the touch display screen is installed inside the frame for user interaction. The processor, storage unit, audio input / output unit, and network communication unit are integrated on the back of the touch display screen and are connected to the touch display screen respectively. The knowledge base management module, which is built into the storage unit, is used to store and manage knowledge data about historical figures. Upload historical figure knowledge to the knowledge base management module, such as relevant knowledge data about historical figure A. This data will correspond to a personal knowledge base, including: Knowledge data: Biography (XXXX-XXXX), major events (events in XXXX corresponding to Person A), military strategies, cultural contributions, etc.; Multimedia data: Historical photos of Person A, scene images of "events in XXXX corresponding to Person A"; Lip shape data: Pre-stored lip movement displacement point data for corresponding Chinese characters and words.

[0023] The digital human presentation module, built into the processor and displayed on the touch screen, is used to generate and drive the digital human image to display content; Facial features, such as facial features, are extracted from historical photos to construct the basic image of a digital human. When displayed, the lip movements will dynamically change in sync with the voice broadcast. By pre-storing the lip movement data of the corresponding text, which contains the coordinates of multiple lip movement displacement points, the lip movement data is matched according to the order of the text fields to be played. The dynamic changes of the lip movements are achieved by controlling the movement of the displacement points, ensuring that the lip movements of the digital human are synchronized with the voice when "speaking".

[0024] The AI ​​interaction module, built into the processor, is used to process user input and extract corresponding response content from the knowledge base management module or obtain corresponding response content by connecting to the cloud server through the network communication unit. It then displays the information and plays audio on the touch screen by calling the digital human presentation module, while simultaneously updating the knowledge base management module.

[0025] The digital human figure is displayed on the touch screen, and the display scene is switched in combination with the question and answer content. For example, when explaining "the event corresponding to person A in XXXX year", the display scene is switched to the corresponding historical scene image.

[0026] When the processor runs the system program and selects the historical figure to be popularized, the digital human presentation module displays the digital human image of figure A on the touch screen, and the audio player plays a welcome message, such as "Hello, I am the digital human of figure A. Welcome to ask me historical questions." It can also list at least one question for the user to choose from.

[0027] The knowledge base management module contains pre-input historical figure knowledge, corresponding character lip-sync data, and corresponding historical figure images. Each historical figure knowledge corresponds to a figure introduction and at least one example. Each historical figure corresponds to a digital human image, which is composed of historical figure photos. Its lip-sync has multiple displacement points, and the lip-sync dynamically changes by moving the displacement points. When the AI ​​interaction module obtains new examples of historical figures, it adds the new examples to the database of that historical figure.

[0028] Upload relevant documents, images, and videos of historical figures to a MySQL database and perform semantic retrieval using the Weaviate database. For example, construct a knowledge graph of figure A, integrate it with a large model, and cover dimensions such as life events, military strategies, and cultural contributions. Perform word segmentation and embedding vectorization on the uploaded text, and store it in the distributed storage service MinIO.

[0029] The audio input / output unit includes a microphone array and an audio player. The microphone array collects ambient human voices and sends them to the processor for analysis. The processor then extracts the words from the human voices and displays them one by one on the touch screen. The AI ​​interaction module analyzes the keywords from the knowledge base management module and extracts corresponding historical figures and their corresponding examples. The audio player then plays the corresponding examples.

[0030] The smart photo frame is fixed to the wall and connected to a power source and network (Wi-Fi or wired). The touch screen, processor, storage unit, microphone array, camera, and audio player are connected through internal wiring to ensure normal communication between the hardware modules.

[0031] The digital human presentation module matches and obtains the corresponding lip-sync data based on the fields to be played, arranges them according to the field order, displays the images of the corresponding historical figures on the touch screen, identifies the lip-sync positions in the images, and dynamically changes the lip-sync based on the lip-sync data.

[0032] The AI ​​interaction module includes a speech recognition module, a natural language understanding unit, and a question-and-answer generation unit. The speech recognition module identifies human voice data collected by the microphone array and converts the voice data into audio text. The natural language understanding unit parses the audio text and the text input through the touch screen to obtain the user's semantics and intent, and displays the audio text and the input text on the touch screen. The question-and-answer generation unit matches corresponding historical figure knowledge and photos within the knowledge base management module based on semantics and intent. The ASR engine converts speech into text, and the BERT model parses the semantics.

[0033] The user first inputs a question via voice or text on the touchscreen, such as "What was the cause of the event in XXXX corresponding to Person A?". A microphone array collects the voice data, and the touchscreen receives the text data. The voice recognition module then converts the collected voice data into text, which is simultaneously displayed on the touchscreen. The natural language understanding unit parses the voice or input text, extracting keywords such as "Person A," "the event in XXXX corresponding to Person A," user intent, and contextual information. Subsequently, the question-and-answer generation unit performs knowledge matching based on the semantic analysis results, prioritizing quick matching of corresponding historical figures' knowledge entries in the local knowledge base management module using keywords. If the time and cause of "Person A and Event" fail to match locally, a request is sent to the cloud server via the network communication unit. The cloud server first searches the cloud knowledge base, then calls the vector database for semantic similarity retrieval. If a match is still not found, a large language model (such as Qwen-7B-Chat) is called to generate an answer, and the newly acquired knowledge is synchronized to the local knowledge base under the knowledge base management module. The question-and-answer generation unit generates related questions based on the current interaction content. For example, after answering "the time of the event", it generates "which specific battles should have occurred for this event", which are displayed on the touch screen and read aloud. Users can click on the question or repeat it by voice to enter the next round of interaction.

[0034] This embodiment has built-in voice output modes for at least one language, such as Mandarin, Minnan dialect, English, and Japanese, and automatically matches the output language according to the user's input language.

[0035] The question-and-answer generation unit matches the content with the knowledge base management module by extracting keywords from audio or text. First, it quickly matches and extracts data from the knowledge base management module using these keywords. If no match is found, it communicates with the cloud server via the network communication unit, sending the corresponding request. The cloud server then performs a semantic similarity search on the cloud knowledge base. If a match is found, the corresponding data is sent back to the knowledge base management module. If the cloud knowledge base still cannot match, it searches the vector database on the cloud server. If a satisfactory result is still not obtained, it uses a large language model to generate the answer.

[0036] The question-and-answer generation unit contains audio and text data. It generates questions associated with the current text data and prepares example data and lip-reading data for each question. The text of the question is displayed on the touch screen and the audio is played back. Users can click on the corresponding question on the touch screen to enter the next stage of the corresponding example display and playback, or repeat the dialogue to the question. The AI ​​interaction module recognizes the keywords in the dialogue, matches the question, and enters the next stage of the corresponding example display and playback.

[0037] The AI ​​interaction module also includes a camera. The camera captures the user's location and image in the current environment, identifies lip movements, determines the user speaking in the frame, identifies and locks onto their voiceprint, and only collects the voice data of that specific voiceprint during the current interaction phase. By using the camera to capture the user's location and image, identify the lip movements and voiceprint of the currently speaking user, and lock onto that user's voiceprint, the module ensures that only their voice data is collected, avoiding interference from multiple users.

[0038] The AI ​​interaction module has a voice output mode for at least one language, which is matched according to the language of the user currently speaking.

[0039] The network communication unit includes a wireless communication module and a wired communication module.

[0040] The specific implementation process is as follows: When a user asks, "In which year did Person A experience the XX event?" Input acquisition: The microphone array captures the user's voice, and the camera simultaneously captures the user's image and lip movements; Speech-to-text: The speech recognition module converts speech into text, such as "In which year did Person A experience XX event?", which is then displayed on the touchscreen. Semantic analysis: The natural language understanding unit extracts keywords "person A", "event XX", and "year" to determine that the user's intent is to query the time of a historical event; Knowledge matching: The question-and-answer generation unit matches the knowledge entry "the event in XXXX year corresponding to person A" in the local knowledge base; Digital Human Demonstration and Broadcast: The digital human presentation module calls up the digital human image of Person A, matches the lip-sync data corresponding to the text "XXXX year event", and controls the dynamic change of the lip-sync displacement point; The audio player synchronously plays the voice message "Person A experienced an event in XXXX year"; Multi-round interactive guidance: The question and answer generation unit generates related questions such as "After the XX event occurred, what measures did Person A take in response to the XX event?", which are displayed on the touch screen and read aloud. Users can click on the question or ask it by voice to enter the next round of interaction.

[0041] When a user asks for content not included in the local knowledge base, such as "What specific measures / data did Person A take in response to Event XX?": If the local knowledge base fails to match, the AI ​​interaction module sends a query request to the cloud server through the network communication unit. After the cloud server failed to find the cloud knowledge base, it called the vector database for semantic retrieval, but still failed to obtain results. Therefore, it called the large language model to generate the answer "When dealing with the XX event, person A adopted XX measures / used XX amount of XX". The cloud server returns the answer to the local knowledge base management module, which automatically adds the content to the "Military Strategy / Life Story" entry in Person A's knowledge base; The system reads the answers aloud using a digital human-like avatar and updates the content displayed on the touchscreen simultaneously.

[0042] The frame of this implementation adopts a lightweight design, which can be fixed to the wall or placed on the table, making it suitable for various scenarios such as home and exhibition hall; The touchscreen display is a 20.5-inch IPS screen with a resolution of 1080×1920 and supports 10-point touch. It is used to display images of historical figures, interactive text content, and dynamic images of digital humans. The processor uses an embedded chip that integrates a computing core and an AI acceleration unit to support the operation of local speech recognition, natural language processing and other algorithms. The storage unit includes local flash memory and expandable storage for storing knowledge base data, dynamic lip-sync data, and system programs; The audio input / output unit consists of a 6-microphone circular array and a high-definition audio player. The microphone array supports far-field voice acquisition up to 2 meters, and the audio player supports high-definition audio output in multiple languages. The network communication unit integrates wireless communication modules, such as Wi-Fi 6 and 4G, and wired communication modules, such as Ethernet interfaces, to enable local and cloud data interaction and support offline / online dual-mode operation. The camera is a high-definition wide-angle camera used to collect data on the user's location, appearance, and lip movements.

[0043] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An AI digital human intelligent photo frame knowledge question-and-answer system, characterized in that, include: A smart photo frame includes a frame, a touch display screen, a processor, a storage unit, an audio input / output unit, and a network communication unit. The frame is placed on a wall, and the touch display screen is installed inside the frame for user interaction. The processor, storage unit, audio input / output unit, and network communication unit are integrated on the back of the touch display screen and are connected to the touch display screen respectively. The knowledge base management module, which is built into the storage unit, is used to store and manage knowledge data about historical figures. The digital human presentation module, built into the processor and displayed on the touch screen, is used to generate and drive the digital human image to display content; The AI ​​interaction module, built into the processor, is used to process user input and extract corresponding response content from the knowledge base management module or obtain corresponding response content by connecting to the cloud server through the network communication unit. It then displays the information and plays audio on the touch screen by calling the digital human presentation module, while simultaneously updating the knowledge base management module.

2. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 1, characterized in that: The knowledge base management module contains pre-input historical figure knowledge, corresponding character lip-sync data, and corresponding historical figure images. Each historical figure knowledge corresponds to a figure introduction and at least one example. Each historical figure corresponds to a digital human image, which is composed of historical figure photos. Its lip-sync has multiple displacement points, and the lip-sync dynamically changes by moving the displacement points. When the AI ​​interaction module obtains new examples of historical figures, it adds the new examples to the database of that historical figure.

3. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 2, characterized in that: The audio input / output unit includes a microphone array and an audio player. The microphone array collects human voices from the surrounding environment and sends them to the processor for analysis to obtain the words in the human voices. These words are then displayed one by one on the touch screen. The AI ​​interaction module analyzes the keywords from the knowledge base management module and extracts corresponding examples of historical figures. The corresponding examples are then played through the audio player.

4. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 3, characterized in that: The digital human presentation module matches and obtains the corresponding lip-sync data based on the fields to be played, arranges them according to the field order, displays the images of the corresponding historical figures on the touch screen, identifies the lip-sync positions in the images, and dynamically changes the lip-sync based on the lip-sync data.

5. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 3, characterized in that: The AI ​​interaction module includes a speech recognition module, a natural language understanding unit, and a question-and-answer generation unit. The speech recognition module identifies human voice data collected by the microphone array and converts the human voice data into audio text. The natural language understanding unit parses the audio text and the text input through the touch screen to obtain the user's semantics and intent, and displays the audio text and the input text on the touch screen. The question-and-answer generation unit matches corresponding historical figure knowledge and historical figure photos within the knowledge base management module based on semantics and intent.

6. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 5, characterized in that: The question-and-answer generation unit matches the content of the knowledge base management module by extracting keywords from the audio or text. First, it quickly matches and extracts data from the knowledge base management module using these keywords. If no match is found, it communicates with the cloud server via the network communication unit, sending the corresponding request. The cloud server then performs a semantic similarity search on the cloud knowledge base. If a match is found, the corresponding data is sent back to the knowledge base management module. If the cloud knowledge base still cannot match, it retrieves data from the cloud server's vector database. If a satisfactory result is still not obtained, it uses a large language model to generate the answer.

7. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 6, characterized in that: The question-and-answer generation unit contains audio data and text data. It generates questions associated with the current text data and prepares example data and lip-reading data corresponding to each question. The text of the question is displayed on the touch screen and the audio is played. Users can click on the corresponding question on the touch screen to enter the next stage of the corresponding example display and playback, or repeat the dialogue of the question. The AI ​​interaction module recognizes the keywords of the dialogue, matches the question, and enters the next stage of the corresponding example display and playback.

8. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 1, characterized in that: The AI ​​interaction module also includes a camera, which captures the user's location and image in the current environment, identifies changes in their lip movements, determines the user speaking in the scene, identifies and locks their voiceprint, and only collects the voice data of that voiceprint during the current interaction phase.

9. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 8, characterized in that: The AI ​​interaction module has at least one language voice output mode, which is matched according to the language of the user currently speaking.

10. The AI ​​digital human intelligent photo frame knowledge question-and-answer system as described in claim 1, characterized in that: The network communication unit includes a wireless communication module and a wired communication module.