Intelligent speech-based image retrieval system
Patent Information
- Authority / Receiving Office
- TW · TW
- Patent Type
- Utility models
- Current Assignee / Owner
- TAIWAN SECOM CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-08-01
Smart Images

Figure TWG2TB001904409_001 
Figure TWG2TB001904409_002 
Figure TWG2TB001904409_003
Abstract
Claims
1. A smart voice and video retrieval system, comprising: a recording module configured to receive an original video captured by at least one camera and store the original video according to at least one channel corresponding to the recording module; an artificial intelligence image recognition module electrically connected to the recording module, configured to sample an image frame of the original video, and recognize the sampled image frame through an artificial intelligence algorithm to generate a corresponding image description text, establish an association between the image description text and the image frame to generate searchable video data, and store the searchable video data in the recording module; and a voice search module electrically connected to the recording module, configured to receive a query voice through a microphone, convert the query voice into query text using voice recognition technology, and then search the searchable video data stored in the recording module based on the query text to obtain a search result, and present the search result through an output display interface.
2. The intelligent voice and video retrieval system as described in claim 1, wherein the artificial intelligence video recognition module is further configured to identify a specific recognition item contained in the sampled video frame through the artificial intelligence algorithm, generate recognition information of the specific recognition item, and establish an association between the recognition information and the video frame to generate the searchable video data.
3. The intelligent voice and image retrieval system as described in claim 2, wherein the voice search module includes a search panel, the search panel including: a voice recognition function area configured to display the microphone's recording status through a dynamic graphic, and to display the query text converted from the query voice; wherein, The microphone is constantly in a listening state and can control the start and end of the voice recognition process through a voice start word and a voice end word.
4. The intelligent voice and image retrieval system as described in claim 3, wherein the voice format of the query voice includes a time interval, a channel identification code, and a search target.
5. The intelligent voice and image retrieval system as described in claim 4, wherein the search panel further includes: a time interval filtering area, configured to automatically select the time interval by automatically inputting a start time and an end time based on the query text; a channel filtering area, configured to automatically select the channel identifier based on the query text; and an event type filtering area, configured to automatically select the search target based on the query text; wherein, The voice search module is further configured to generate a filtering condition based on the time interval, channel identifier, and search target selected on the search panel.
6. The intelligent voice and video retrieval system as described in claim 5, wherein each of the filtering conditions in the time interval filtering area, the channel filtering area and the event type filtering area of the search panel can also be manually input or selected through an input device.
7. The intelligent voice and image retrieval system as described in claim 5, wherein the voice search module is configured to execute an event search procedure or an object search procedure based on the filtering criteria; wherein, When the selected search target is an event type, the event search procedure is triggered to compare the query text with the associated image description text in the searchable video data to filter out matching image frames; and when the selected search target is an object type, the object search procedure is triggered to compare one of the objects in the query text that matches the specific identification item with the identification information of the specific identification item associated with the searchable video data to filter out matching image frames.
8. The intelligent voice and video retrieval system as described in claim 7, wherein the artificial intelligence video recognition module is further configured to instantly recognize the object in the video frame corresponding to the specific recognition item, generate a bounding box for the recognized object, draw the bounding box on the video frame, and establish an association between the bounding box and the video frame to generate the searchable video data.
9. The intelligent voice and image retrieval system as described in claim 8, wherein the voice search module is configured to execute the object search procedure according to the filtering conditions, the object search procedure comparing the object that matches the specific identification item with the bounding box associated in the searchable video data to filter out the matching image.
10. The intelligent voice and image retrieval system as described in claim 1, wherein the artificial intelligence image recognition module is further configured to identify at least one human individual in the sampled image frame through the artificial intelligence algorithm, and generate a bounding box and a personal identification code for the identified human individual, draw the bounding box on the image frame, and establish an association between the personal identification code and the image frame to generate the searchable video data, wherein in the searchable video data, the same human individual corresponds to the same personal identification code.
11. The intelligent voice and video retrieval system as described in claim 6, further comprising: a statistical analysis module electrically connected to the recording module, configured to execute a data processing program to generate a statistical analysis report; wherein, The data processing program includes: receiving and parsing the searchable video data stored by the video recording module to obtain a bounding box and a personal identification code corresponding to each human individual in each video frame; extracting time information corresponding to the video frame based on the bounding box and the personal identification code, and calculating coordinate information of the bounding box; tracking the changes in the coordinate information of the bounding box in consecutive video frames through a tracking algorithm to maintain the consistency of the personal identification code of the same human individual in different video frames; establishing a correlation between the bounding box, the personal identification code, the time information, and the coordinate information in each video frame to form a structured dataset; and performing statistical analysis calculations based on the structured dataset.
12. The intelligent voice and image retrieval system as described in claim 11, wherein the data processing program includes generating at least one chart based on the structured dataset using a statistical analysis tool, and outputting the at least one chart to the output display interface, wherein, The at least one chart is selected from one or more of the following charts: a trend chart of total number of people over time, a trend chart of flow density over time, a bar chart of average number of people, and a heat map of flow density.