Artificial intelligence three-dimensional digital interactive display device

By designing an artificial intelligence three-dimensional digital interactive display device, using speech recognition, vectorized data retrieval and multimodal display technology, the problems of rigid content display and poor interactive experience of traditional devices are solved, and customized, intuitive and efficient data display is achieved.

CN119961647AInactive Publication Date: 2025-05-09MUSTARD SEED DATA (GUANGZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510047007.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The content display of traditional three-dimensional digital interactive display devices is rigid, lacking personalization and flexibility, and cannot provide customized content display according to the needs of different users. The effective interactive experience provided is poor, and it fails to make full use of the structured and unstructured data accumulated by the enterprise.

Method used

An artificial intelligence three-dimensional digital interactive display device is designed, including a speech recognition module, a search module, a database module and a display module. The voice recognition module captures user voice commands through directional sound pickup technology, the database module quickly searches the data vectors, and the display module provides interactive display through the touch screen and the holographic unit.

Benefits of technology

It realizes customized data retrieval and display based on user voice commands, provides an intuitive and interactive user experience, improves the efficiency and accuracy of data retrieval, and makes full use of the structured and unstructured data of the enterprise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961647A_ABST
    Figure CN119961647A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence three-dimensional digital interactive display device, and relates to the technical field of data display, the artificial intelligence three-dimensional digital interactive display device comprises a voice recognition module, a search module, a database module and a display module, the voice recognition module is used for capturing a voice instruction of a user and comprises an activation unit, a noise analysis unit, a directional pickup unit and a voice processing unit; the activation unit is started in a preset activation mode, and after the activation unit is started, the noise analysis unit and the directional pickup unit are activated; and the noise analysis unit performs noise feature analysis on the acquired environment audio data to generate a noise reduction model. Through the directional pickup technology, target voice signals can be effectively captured, background noise in the non-target direction can be suppressed, the voice recognition capability in a complex environment is improved, and in the voice instruction processing stage, a dynamic noise model generation function is introduced to adapt to environment changes in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data display, and in particular to an artificial intelligence three-dimensional digital interactive display device. Background Art

[0002] With the rapid development of artificial intelligence, big data and visualization technology, three-dimensional digital interactive display devices have gradually become important tools in the fields of industrial design, education and training, medical image analysis, architectural modeling, etc. Such devices usually help users analyze and understand complex data more intuitively by integrating voice recognition, data processing and multi-dimensional display technology.

[0003] After searching, a Chinese patent (publication number: CN113332692A) discloses a digital immersive interactive device based on artificial intelligence, which includes a cockpit, a cabin door at the rear side of the cockpit, a seat device at the rear part of the bottom of the cockpit, two side walls of the cockpit connected to a first rotating shaft, the first rotating shafts are connected by a sprocket transmission assembly, a drawing board is provided on the outside of the chain of the sprocket transmission assembly, annular slide rails are provided on the left and right sides of the seat device, a first slider is provided on the inner side of the annular slide rail, a handle is provided on the first slider, the handle is connected to the corresponding first rotating shaft through a connecting device, and an electronic display is provided on the front side of the interior of the cockpit.

[0004] Traditional display devices display content rigidly, lack personalization and flexibility, cannot provide customized content display according to the needs of different users, and provide poor effective interactive experience, cannot effectively interact with the displayed content, and fail to fully utilize the structured and unstructured data accumulated by the enterprise. Therefore, the present invention provides an artificial intelligence three-dimensional digital interactive display device. Summary of the invention

[0005] The purpose of the present invention is to provide an artificial intelligence three-dimensional digital interactive display device to solve the problems mentioned in the above background technology.

[0006] The present invention can be implemented by the following technical solutions: an artificial intelligence three-dimensional digital interactive display device, including a speech recognition module, a search module, a database module and a display module;

[0007] The speech recognition module is used to capture the user's voice commands and convert them into executable search query information;

[0008] The database module is used to store various types of data accumulated by users, and the database module vectorizes various types of data to obtain data vectors of various data, and converts text, images, videos and design files into vector representations in high-dimensional space, thereby achieving rapid retrieval and matching;

[0009] The search module is connected to the database module and recalls the most relevant data from the database module based on the search query information converted by the speech recognition module;

[0010] The display module includes an interactive unit and a holographic unit for displaying the data recalled from the database;

[0011] The interactive unit uses a touch screen, and the user can perform interactive operations on the displayed content through the touch screen, such as rotating the model or enlarging the drawing;

[0012] The holographic unit displays three-dimensional models and designs in a stereoscopic form, and users can directly view and evaluate the design content.

[0013] A further technical improvement of the present invention is that: the speech recognition module includes an activation unit, a noise analysis unit, a directional sound pickup unit and a speech processing unit;

[0014] The activation unit is started by a preset activation method, and after being started, activates the noise analysis unit and the directional sound pickup unit, which specifically includes a touch control part and a voice activation part;

[0015] The touch part is for the user to start the voice activation part by physically touching or touching a specific structure of the device;

[0016] After the voice activation part is started, perform the following operations:

[0017] Signal acquisition, collecting real-time audio signals through low-power microphones;

[0018] Feature extraction, using the Mel frequency spectrum features (MFCC) of the speech signal as input and identifying the wake-up word through the convolutional neural network (CNN) model;

[0019] Logical judgment: when the activation unit detects that the recognized wake-up word is the same as the preset wake-up word, the activation unit starts, and collects environmental audio data when the noise analysis unit and the directional sound pickup unit are activated;

[0020] The noise analysis unit performs noise feature analysis on the collected ambient audio data to generate a noise reduction model for subsequent directional sound pickup and speech processing, and the specific steps include:

[0021] Signal processing: Use Fast Fourier Transform (FFT) to convert audio signals into frequency domain, extract spectrum features, and calculate the time domain and spectrum features of noise;

[0022] Feature classification: Use machine learning models to classify noise types;

[0023] Noise model generation: Generate a dynamic noise model based on the feature classification results, including frequency range and noise reduction intensity parameters;

[0024] The directional sound pickup unit determines the target speech direction through microphone array technology, enhances the target speech signal, and suppresses noise in non-target directions;

[0025] Microphone array technology uses a circular or linear microphone array, where each microphone captures sound signals at different locations. It determines the direction of the sound source by calculating the time difference between the target speech reaching each microphone, and further corrects the direction by analyzing the difference in sound pressure intensity at different locations in the microphone array.

[0026] After confirming the direction of the sound source, the directional sound pickup unit uses a beamforming algorithm to enhance the target direction signal, and uses the noise reduction model generated by the noise analysis unit to further filter the background noise signal in the direction of the sound source;

[0027] The voice processing unit converts the target voice signal output by the directional sound pickup unit into search query information.

[0028] A further technical improvement of the present invention is that: after receiving the target voice signal output by the directional sound pickup unit, the voice processing unit calculates the signal decibel of the target voice signal and compares the signal decibel of the target voice signal with a preset decibel threshold;

[0029] If the signal decibel of the target voice signal is lower than the decibel threshold, the voice processing unit abandons converting the target voice signal this time and waits for the next target voice signal transmission;

[0030] If the signal decibel of the target voice signal is not lower than the decibel threshold, the voiceprint recognition process will be entered.

[0031] A further technical improvement of the present invention is that the process of voiceprint recognition includes:

[0032] S1, select the voiceprint of the target voice signal before S time as the reference voiceprint of the current voice command;

[0033] S2. Compare the target voice signal of the remaining part of the voice command with the reference voiceprint.

[0034] S3. After the voiceprint recognition is passed, the voice processing unit converts the target voice signal into search query information.

[0035] A further technical improvement of the present invention is that the step of the speech processing unit outputting the search query information comprises:

[0036] Q1. The speech processing unit converts the target speech signal into a text instruction through a speech recognition model. In this embodiment, the DeepSpeech model is used. For example, if the user says "search for car design drawings", the corresponding text instruction "search for car design drawings" is output;

[0037] Q2, the speech processing unit uses the natural language processing model to extract key entities and data types in the text instructions;

[0038] Q3. Identify user operation goals through deep learning models;

[0039] Q4. Use a pre-trained language model to convert text instructions into high-dimensional vectors.

[0040] A further technical improvement of the present invention is that the step of the search module recalling the most relevant data includes:

[0041] Z1, receiving search query information output by the speech processing unit, including key entities, data types, intents, and high-dimensional vectors;

[0042] Z2. Filter the database, including:

[0043] Filter data types: according to the data type in the search query information;

[0044] Key entity matching: further filter the records in the database module based on the key entities in the search query information;

[0045] Z3, perform vector similarity matching;

[0046] The vector matching tool FAI SS is used to calculate the similarity between the high-dimensional vector output by the speech processing unit and each data vector in the database module. The formula used is: In the formula, X is the similarity between the high-dimensional vector and the data vector, is the high-dimensional vector output by the speech processing unit, is the data vector in the database module;

[0047] And arrange each data vector in descending order according to similarity to obtain the matching result;

[0048] Z4. Package the matching results into a unified format, including data ID, type, similarity score and specific content path, and return the matching results to the display module;

[0049] The display module displays each data in the matching result in a carousel, and determines whether to pause the carousel based on whether there is an operation action.

[0050] A further technical improvement of the present invention is that: when storing various types of data, the database module presets an introduction text for each type of data, and the database module performs speech synthesis processing on each introduction text;

[0051] When the display module displays the recalled data, the corresponding introduction text will be played;

[0052] When displaying important design documents or key production data, users can quickly understand the file content without having to manually view each file. The application of automatic broadcast technology not only improves the efficiency of information transmission, but also enhances the user experience.

[0053] A further technical improvement of the present invention is that: the database module includes a classification unit, and the classification unit is used to classify each data according to the type of picture, video, two-dimensional drawing and three-dimensional model;

[0054] And the classification unit groups various types of data according to the corresponding processes.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] The present invention can effectively capture the target voice signal and suppress the background noise in the non-target direction through directional sound pickup technology, thereby improving the voice recognition capability in complex environments. In addition, in the voice command processing stage, a dynamic noise model generation function is introduced to adapt to environmental changes in real time, further enhancing the accuracy and robustness of recognition.

[0057] At the same time, by using deep learning models to extract features from text, images, videos, and 3D model data, all data is uniformly mapped to a high-dimensional vector space, thereby achieving efficient matching of cross-modal data, and using efficient vector retrieval tools to quickly filter out multimodal data that is most relevant to the user's voice query;

[0058] On the other hand, the present invention quickly triggers the retrieval and display functions through voice commands, and the display module responds to the user's touch operations in real time, supporting efficient interaction combined with multiple modalities. The entire workflow can improve the efficiency and accuracy of data retrieval while providing an intuitive and interactive user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.

[0060] Figure 1 is a system block diagram of the present invention;

[0061] Figure 2 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0062] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0063] See also Figure 1-2 As shown, the present invention provides an artificial intelligence three-dimensional digital interactive display device, including a speech recognition module, a search module, a database module and a display module;

[0064] The speech recognition module is used to capture the user's voice commands and convert them into executable search query information;

[0065] The speech recognition module includes an activation unit, a noise analysis unit, a directional sound pickup unit and a speech processing unit;

[0066] The activation unit is started by a preset activation method, and after being started, activates the noise analysis unit and the directional sound pickup unit, which specifically includes a touch control part and a voice activation part;

[0067] The touch part is for the user to start the voice activation part by physically touching or touching a specific structure of the device;

[0068] In this embodiment, the touch control part adopts a microphone button. When in use, the voice activation part is activated by touching the microphone button;

[0069] After the voice activation part is started, perform the following operations:

[0070] Signal acquisition, collecting real-time audio signals through low-power microphones;

[0071] Feature extraction: Use the Mel frequency spectrum feature (MFCC) of the speech signal as input and use the convolutional neural network (CNN) model to identify the wake-up word. When extracting features, determine whether it exceeds 30 seconds. If the wake-up word is not recognized after 30 seconds, the speech recognition is terminated.

[0072] Logical judgment: when the activation unit detects that the recognized wake-up word is the same as the preset wake-up word, the activation unit starts, and collects environmental audio data when the noise analysis unit and the directional sound pickup unit are activated;

[0073] The noise analysis unit performs noise feature analysis on the collected ambient audio data to generate a noise reduction model for subsequent directional sound pickup and speech processing, and the specific steps include:

[0074] Signal processing: Use Fast Fourier Transform (FFT) to convert audio signals into frequency domain, extract spectrum features, and calculate the time domain and spectrum features of noise;

[0075] Feature classification: Use machine learning models to classify noise types;

[0076] Noise model generation: Generate a dynamic noise model based on the feature classification results, including frequency range and noise reduction intensity parameters;

[0077] The directional sound pickup unit determines the target speech direction through microphone array technology, enhances the target speech signal, and suppresses noise in non-target directions;

[0078] Microphone array technology uses a circular or linear microphone array, where each microphone captures sound signals at different locations. It determines the direction of the sound source by calculating the time difference between the target speech reaching each microphone, and further corrects the direction by analyzing the difference in sound pressure intensity at different locations in the microphone array.

[0079] After confirming the direction of the sound source, the directional sound pickup unit uses a beamforming algorithm to enhance the target direction signal, and uses the noise reduction model generated by the noise analysis unit to further filter the background noise signal in the direction of the sound source;

[0080] After receiving the target voice signal output by the directional sound pickup unit, the voice processing unit calculates the signal decibel of the target voice signal and compares the signal decibel of the target voice signal with a preset decibel threshold;

[0081] If the signal decibel of the target voice signal is lower than the decibel threshold, the voice processing unit abandons converting the target voice signal this time and waits for the next target voice signal transmission;

[0082] If the signal decibel of the target voice signal is not lower than the decibel threshold, the voiceprint recognition process will be entered;

[0083] The voiceprint recognition process includes:

[0084] S1, select the voiceprint of the target voice signal before S time as the reference voiceprint of the current voice command;

[0085] S2, comparing the target voice signal of the remaining part of the voice command with the reference voiceprint;

[0086] In this embodiment, cosine similarity is used for comparison. If the similarity between the two is higher than a preset similarity threshold, it is determined that the target voice signal is a voice command issued by the same user.

[0087] The formula used is In the formula, S i is the similarity, A is the voiceprint of the current target speech signal, and B is the reference voiceprint;

[0088] And the similarity threshold is set to 80%.i ≥80%, it is determined to be the same voiceprint. i When the rate is less than 80%, the voice command is discarded;

[0089] After the voiceprint recognition is passed, the voice processing unit converts the target voice signal into search query information;

[0090] The speech processing unit converts the target speech signal output by the directional sound pickup unit into search query information, and the steps include:

[0091] Q1. The speech processing unit converts the target speech signal into a text instruction through a speech recognition model. In this embodiment, the DeepSpeech model is used. For example, if the user says "search for car design drawings", the corresponding text instruction "search for car design drawings" is output;

[0092] Q2, the speech processing unit uses the natural language processing model to extract key entities and data types in the text instructions;

[0093] For example, in the text instruction "search for car design drawings", the key entity "car" and the data type "design drawings" are extracted;

[0094] Q3. Identify user action goals through deep learning models, such as extracting the intent “search” in “search for car design drawings”;

[0095] Q4. Use the pre-trained language model to convert text instructions into high-dimensional vectors. For example, when combining "car", "design drawing", and "search", the high-dimensional vector [0.45, -0.32, 0.67, ...] is output;

[0096] The database module is used to store various types of data accumulated by users, and the database module vectorizes various types of data to obtain data vectors of various data, and converts text, images, videos and design files into vector representations in high-dimensional space, thereby achieving rapid retrieval and matching;

[0097] For example, when extracting text data, we use word embedding technology to map the text content to the vector space and convert the sentence "design drawings" into a vector [0.56, 0.34, -0.22, ...];

[0098] When extracting image data, use the ResNet deep learning model to extract the feature vector of the image, for example, convert the image into a vector [0.89, -0.76, 1.22, ...];

[0099] When extracting video data, use LSTM to extract time series features, such as converting the video into [0.45, 0.67, 0.89,...];

[0100] When extracting 3D model data, extract its geometric features, including vertices, normal vectors, and surface textures, and generate vector representations, such as converting the model point cloud into [1.2, 0.3, -0.8, ...];

[0101] And the database module includes a classification unit, which is used to classify each data according to the type of picture, video, two-dimensional drawing and three-dimensional model;

[0102] And the classification unit groups various types of data according to the corresponding processes. For example, when grouping according to the production process, the design drawings will be classified into the "design file" group, and the production video will be classified into the "production process" group;

[0103] The search module is connected to the database module and recalls the most relevant data from the database module based on the search query information converted by the speech recognition module;

[0104] The search module recalls the most relevant data steps, including:

[0105] Z1, receiving search query information output by the speech processing unit, including key entities, data types, intents, and high-dimensional vectors;

[0106] Z2. Filter the database, including:

[0107] Filter data types: According to the data type in the search query information, in this embodiment, it is "design drawings", filter the corresponding data type records in the database module;

[0108] Key entity matching: further filter the records in the database module according to the key entity in the search query information, which is "car" in this embodiment;

[0109] Z3, perform vector similarity matching;

[0110] The vector matching tool FAI SS is used to calculate the similarity between the high-dimensional vector output by the speech processing unit and each data vector in the database module. The formula used is: In the formula, X is the similarity between the high-dimensional vector and the data vector, is the high-dimensional vector output by the speech processing unit, is the data vector in the database module;

[0111] And arrange each data vector in descending order according to similarity to obtain the matching result;

[0112] Z4. Package the matching results into a unified format, including data ID, type, similarity score and specific content path, and return the matching results to the display module;

[0113] The display module performs a carousel on each data in the matching result, and determines whether to pause the carousel based on whether there is an operation action. For example, when the user clicks the touch screen, the carousel is paused, and after the corresponding data is confirmed, the carousel is stopped;

[0114] The display module includes an interactive unit and a holographic unit for displaying the data recalled from the database;

[0115] The interactive unit uses a touch screen, and the user can perform interactive operations on the displayed content through the touch screen, such as rotating the model or enlarging the drawing;

[0116] The holographic unit displays the three-dimensional model and design in a three-dimensional form, and the user can directly view and evaluate the design content. In this embodiment, the holographic unit uses holographic projection technology to generate a three-dimensional image in space. Specifically, corresponding model equipment can be used in the existing technology as needed.

[0117] Example 2

[0118] Compared with Example 1, the display module in Example 2 replaces the holographic unit with an AR unit, uses augmented reality technology, and superimposes the three-dimensional model and design file into the user's actual field of view through AR applications on smart glasses or mobile devices to achieve interactive display. This solution can use existing smart phones and tablets, reducing the need for additional hardware.

[0119] Example 3

[0120] Compared with Example 1, the display module in Example 3 replaces the holographic unit with a VR unit. Through a virtual reality helmet or VR glasses, the user can enter a completely virtual environment and interact immersively with three-dimensional models and design files.

[0121] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. An artificial intelligence three-dimensional digital interactive display device, comprising a speech recognition module, a search module, a database module and a display module, characterized in that: The speech recognition module is used to capture the user's voice command, and includes an activation unit, a noise analysis unit, a directional sound pickup unit and a speech processing unit; The activation unit is started by a preset activation method, and after being started, the noise analysis unit and the directional sound pickup unit are activated; The noise analysis unit performs noise feature analysis on the collected ambient audio data to generate a noise reduction model; The directional sound pickup unit determines the target speech direction through microphone array technology, enhances the target speech signal, and suppresses noise in non-target directions; After confirming the direction of the sound source, the directional sound pickup unit uses a beamforming algorithm to enhance the target direction signal, and uses the noise reduction model generated by the noise analysis unit to further filter the background noise signal in the direction of the sound source; The speech processing unit converts the target speech signal output by the directional sound pickup unit into search query information; The database module is used to store various types of data accumulated by users, and the database module vectorizes various types of data to obtain data vectors of various data; The search module is connected to the database module and recalls the most relevant data from the database module based on the search query information converted by the speech recognition module; The display module is used to display the data recalled from the database.

2. The artificial intelligence three-dimensional digital interactive display device according to claim 1, characterized in that: The display module includes an interactive unit and a holographic unit; The interactive unit uses a touch screen, and the user performs interactive operations on the displayed content through the touch screen; The holographic unit displays three-dimensional models and designs in stereoscopic form.

3. The artificial intelligence three-dimensional digital interactive display device according to claim 1, characterized in that: After receiving the target voice signal output by the directional sound pickup unit, the voice processing unit calculates the signal decibel of the target voice signal and compares the signal decibel of the target voice signal with a preset decibel threshold; If the signal decibel of the target voice signal is lower than the decibel threshold, the voice processing unit abandons converting the target voice signal this time and waits for the next target voice signal transmission; If the signal decibel of the target voice signal is not lower than the decibel threshold, the voiceprint recognition process will be entered.

4. The artificial intelligence three-dimensional digital interactive display device according to claim 3, characterized in that: The voiceprint recognition process includes: S1, select the voiceprint of the target voice signal before S time as the reference voiceprint of the current voice command; S2, comparing the target voice signal of the remaining part of the voice command with the reference voiceprint; S3. After the voiceprint recognition is passed, the voice processing unit converts the target voice signal into search query information.

5. The artificial intelligence three-dimensional digital interactive display device according to claim 1, characterized in that: The step of the speech processing unit outputting the search query information comprises: Q1, the speech processing unit converts the target speech signal into text instructions through the speech recognition model; Q2, the speech processing unit uses the natural language processing model to extract key entities and data types in the text instructions; Q3. Identify user operation goals through deep learning models; Q4. Use a pre-trained language model to convert text instructions into high-dimensional vectors.

6. The artificial intelligence three-dimensional digital interactive display device according to claim 5, characterized in that: The search module recalls the most relevant data steps, including: Z1, receiving search query information output by the speech processing unit, including key entities, data types, intents, and high-dimensional vectors; Z2. Filter the database, including: Filter data types: according to the data type in the search query information; Key entity matching: based on the key entities in the search query information; Z3, perform vector similarity matching, and arrange each data vector in descending order of similarity to obtain a matching result; Z4. Package the matching results into a unified format and return the matching results to the display module.

7. The artificial intelligence three-dimensional digital interactive display device according to claim 6, characterized in that: The display module displays each data in the matching result in a carousel, and determines whether to pause the carousel based on whether there is an operation action.

8. The artificial intelligence three-dimensional digital interactive display device according to claim 1, characterized in that: When storing various types of data, the database module presets an introduction text for each type of data, and the database module performs speech synthesis processing on each introduction text; When the display module displays the recalled data, the corresponding introduction text will be played.

9. The artificial intelligence three-dimensional digital interactive display device according to claim 1, characterized in that: The database module includes a classification unit, which is used to classify each data according to the type of picture, video, two-dimensional drawing and three-dimensional model; And the classification unit groups various types of data according to the corresponding processes.

Citation Information

Patent Citations

  • Digital immersive interaction device based on artificial intelligence

    CN113332692A

  • Pick-up method and system based on microphone array

    CN106782585A

  • Platform for intelligent interaction of popular science knowledge

    CN117540028A