Intelligent data query method and system based on big data

By generating guide data and using a large language model to understand natural language queries, the problem of insufficient accuracy of existing systems in multimodal data processing is solved, and fast and accurate data retrieval and result generation are achieved.

CN120429479BActive Publication Date: 2025-09-26HANGZHOU ZHIYOU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510933005.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-26
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Existing intelligent query systems find it difficult to effectively understand users' natural language query intentions, especially when processing multimodal data. The query results are not accurate and practical enough. Traditional data processing methods are inefficient and cannot meet the needs of quickly obtaining information.

Method used

By reading multimodal data, generating metadata descriptions and content summaries, using a large language model to generate guidance data, storing it in memory and updating it periodically, receiving natural language query requirements, generating matching tag sets, querying matching multimodal data rows, and generating query result reports.

Benefits of technology

It improves the data retrieval speed and the accuracy of query results, can quickly obtain the required information, integrate data from different sources and formats, and generate intuitive query result reports, which is suitable for the rapid retrieval of large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429479B_ABST
    Figure CN120429479B_ABST
Patent Text Reader

Abstract

Multiple embodiments of this specification relate to the field of information technology, and specifically to a method and system for intelligent data query based on big data. The method includes the following steps: reading multimodal data, establishing multimodal data rows according to the source and acquisition time; generating metadata description, content summary, preset scene tag set, and preset scene effect description for each multimodal data row according to the multimodal file description and multimodal file metadata; generating guide data for each multimodal data row; storing the guide data in a pre-opened memory space; receiving a query requirement expressed in natural language, and using a pre-connected large language model to generate a tag set that matches the query requirement; querying the matching guide data according to the tag set to obtain matching multimodal data; and generating a query result report according to the metadata description, content summary, and preset application scenario effect description of the matching multimodal data row.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Multiple embodiments of this specification relate to the field of information technology, and specifically to an intelligent data query method and system based on big data. Background Art

[0002] With the rapid development of information technology, the amount of data is growing explosively. These data include but are not limited to various data types such as text, audio, and video. In the context of big data, how to efficiently store, manage, and query these massive amounts of multimodal data has become an important issue. Traditional data processing methods can often only process a single type of data, lack the ability to support multiple types of data simultaneously, and are inefficient when faced with large-scale data, making it difficult to meet users' needs for quickly obtaining the required information. Especially in the current field of information retrieval, users expect to be able to express their queries in natural language and quickly receive accurate results. However, most existing intelligent query systems are unable to fully understand users' natural language query intent, especially when it comes to complex multimodal data. The accuracy and practicality of their query results still need to be improved. To this end, new data query technologies need to be studied. Summary of the Invention

[0003] Multiple embodiments of this specification describe an intelligent data query method and system based on big data.

[0004] In a first aspect, the embodiments of this specification provide a method for intelligent data query based on big data, comprising the steps of:

[0005] Read multimodal data and create a multimodal data row based on the source and acquisition time. The multimodal data row includes a row ID, a multimodal file name, a multimodal file address, a multimodal file description, and multimodal file metadata.

[0006] Generating a metadata description, a content summary, a preset scene tag set, and a preset scene effect description for each multimodal data row according to the multimodal file description and the multimodal file metadata;

[0007] Generate guidance data for each multi-mode data row, the guidance data including a sequence number, a starting storage address of the multi-mode data row, a metadata description, a content summary, a preset application scenario tag set, and a preset application scenario effect description;

[0008] The guidance data is stored in a pre-opened memory space and is persistently stored according to a preset period. When the multi-mode data is updated, the multi-mode data row and the guidance data are triggered to be updated;

[0009] Receive a query requirement expressed in natural language, and use a pre-connected large language model to generate a tag set that matches the query requirement;

[0010] querying matching guide data according to the tag set, obtaining matching multimodal data rows according to the matching guide data, and obtaining matching multimodal data according to the multimodal data rows;

[0011] Generate a query result report based on the metadata description, content summary, and preset application scenario effect description of the matched multi-modal data rows.

[0012] In a second aspect, the embodiments of this specification provide an intelligent data query system based on big data, including:

[0013] A reading module reads multimodal data and creates a multimodal data row according to the source and acquisition time. The multimodal data row includes a row ID, a multimodal file name, a multimodal file address, a multimodal file description, and multimodal file metadata.

[0014] A first generation module generates a metadata description, a content summary, a preset scene tag set, and a preset scene effect description for each multimodal data row based on the multimodal file description and the multimodal file metadata;

[0015] a second generating module, generating guide data for each multi-mode data row, wherein the guide data includes a sequence number, a starting storage address of the multi-mode data row, a metadata description, a content summary, a preset application scenario tag set, and a preset application scenario effect description;

[0016] A guidance module stores the guidance data in a pre-opened memory space and stores it persistently according to a preset period. When the multi-mode data is updated, it triggers the update of the multi-mode data row and the guidance data.

[0017] A third generation module receives a query requirement expressed in natural language and uses a pre-connected large language model to generate a tag set that matches the query requirement;

[0018] a query module, querying matching guide data according to the tag set, obtaining matching multimodal data rows according to the matching guide data, and obtaining matching multimodal data according to the multimodal data rows;

[0019] The result module generates a query result report based on the metadata description, content summary, and preset application scenario effect description of the matched multi-modal data rows.

[0020] In a third aspect, embodiments of this specification provide an electronic device, including a processor and a memory;

[0021] The processor is connected to the memory;

[0022] The memory is used to store executable program code;

[0023] The processor reads the executable program code stored in the memory to run a program corresponding to the executable program code, so as to execute the method described in any one of the above aspects.

[0024] In a fourth aspect, an embodiment of this specification provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method described in any one of the above aspects is implemented.

[0025] In a fifth aspect, embodiments of this specification provide a computer program product, including a computer program, which implements the method described in any of the above aspects when executed by a processor.

[0026] The beneficial effects of the technical solutions provided by some embodiments of this specification include at least:

[0027] In multiple embodiments of this specification, the provided big data-based intelligent data query method and system effectively improves the speed of data retrieval by performing structured processing on multimodal data (including text, audio, video, etc.) and generating guide data. The required information can be quickly obtained without accessing the original files one by one, which greatly improves the response time. A large language model is used to deeply analyze natural language query requirements, and semantic matching is performed in combination with a preset application scenario tag set to ensure the high relevance and accuracy of the query results. It can effectively integrate and analyze data from different sources and formats, breaking the limitations of traditional single-type data processing. By generating a preliminary query result report, including basic information description, summary description, and effect description, complex big data analysis results become intuitive and easy to understand. It is suitable for performing preliminary and rapid retrieval of large amounts of data, and at the same time obtaining a concise query result report to provide rapid support for subsequent decision-making.

[0028] Other features and advantages of the various embodiments of this specification will be further disclosed in the following detailed description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0030] Figure 1 This is a schematic diagram of the intelligent data query method and system application scenario provided in the embodiments of this specification.

[0031] Figure 2This is a flow chart of the intelligent data query method provided in the embodiments of this specification.

[0032] Figure 3 This is a schematic diagram of the guidance data provided in the embodiments of this specification.

[0033] Figure 4 This is a schematic diagram of the intelligent data query system provided in the embodiments of this specification.

[0034] Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0035] The following is an explanation and description of the technical solutions of the embodiments of this specification in conjunction with the drawings of the embodiments of this specification. However, the following embodiments are only preferred embodiments of this specification and are not exhaustive. Based on the embodiments in the implementation mode, other embodiments obtained by those skilled in the art without making any creative work are all within the scope of protection of this specification.

[0036] Throughout this specification, the claims, and the accompanying drawings, the terms "first," "second," "third," and the like are used to distinguish between different items, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may include other steps or elements inherent to the process, method, product, or apparatus.

[0037] In the following description, terms such as "inside", "outside", "up", "down", "left", "right", etc. that indicate directions or positional relationships are only used to facilitate the description of the embodiments and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, they should not be understood as limitations on this specification.

[0038] The data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions.

[0039] Before introducing the technical solution recorded in this specification, the application scenarios of the technical solution and related technologies are introduced.

[0040] With the development of informatization, the rapid growth of data and the diversification of data types have brought unprecedented challenges to the field of information retrieval. Traditional data retrieval technology is mainly designed for structured data. When faced with massive, heterogeneous data sets (such as text, audio, video, etc.), its efficiency and accuracy are significantly reduced. Specifically, first of all, the fusion processing of multimodal data has become a major problem. Multimedia content has been widely used, and data is no longer limited to a single text form. It also includes a large amount of unstructured data such as audio and video. How to effectively integrate these data in different formats and extract valuable information from them is one of the key issues that current retrieval technology urgently needs to solve.

[0041] The growth of data size has led to an intensified conflict between retrieval speed and resource consumption. Traditional indexing methods often consume significant computing resources and time when processing large datasets, which is unacceptable for applications with high real-time requirements. Furthermore, understanding user query intent and matching accuracy still need to be improved.

[0042] For this purpose, please see the attached Figure 1 This specification provides a new retrieval technology that effectively processes multimodal data, rapidly responds to queries, deeply understands user intent, and automatically generates high-quality content summaries. This not only helps improve data utilization efficiency but also provides users with more convenient and accurate data query services.

[0043] First, this manual provides an intelligent data query method based on big data. Figure 2 , including the steps of:

[0044] Step S1) reads multimodal data 11 and creates a multimodal data row according to the source and acquisition time. The multimodal data row includes a row ID, a multimodal file name, a multimodal file address, a multimodal file description, and multimodal file metadata.

[0045] The multimodal data 11 includes text data, audio data and video data.

[0046] The steps to create multi-mode data rows based on source and acquisition time include:

[0047] Establishing a multimodal data table, wherein the columns of the multimodal data table include an ID column, a text table column, a text address column, a text description column, a text metadata column, an audio table column, an audio address column, an audio description column, an audio metadata column, a video table column, a video address column, a video description column, and a video metadata column;

[0048] The multimode data 11 are grouped according to the source, and the multimode data 11 are truncated according to the generation time according to the preset truncation period. The multimode data 11 of the same group in the same truncation period are regarded as a subgroup, and one subgroup corresponds to one multimode data row;

[0049] Identify the text in the sub-grouped audio data and record it as audio text;

[0050] Identify the sound recognition text of the audio track in the video data, identify the picture of the video, and obtain the picture description text, wherein the sound recognition text and the picture description text constitute the video text;

[0051] Extracting text summaries of the text data, audio text, and video text in the subgroups;

[0052] Extracting the file names, file addresses and metadata of the text data, audio data and video data respectively;

[0053] Add a new row to the multi-mode data table, assign a unique row ID to the new row and fill it in the ID column;

[0054] Fill the file names, file addresses, text summaries and metadata corresponding to the text data, audio data and video data into the corresponding columns.

[0055] Build a multimodal data table to store data from different modalities. This table contains multiple columns, including: an ID column that uniquely identifies each record; text columns that include a text table column, a text address column, a text description column, and a text metadata column; audio columns that include an audio table column, an audio address column, an audio description column, and an audio metadata column; and video columns that include a video table column, a video address column, a video description column, and a video metadata column.

[0056] The collected multi-modal data 11 are preliminarily grouped according to their sources, and each group of data is sliced ​​in the time dimension according to a preset time cut-off period (such as hourly, daily, etc.). Specifically:

[0057] All data originating from the same device or system are grouped together;

[0058] Divide each set of data according to the set truncation period (for example, every hour);

[0059] All multimode data 11 within the same truncation period and from the same source form a subgroup;

[0060] Each subgroup corresponds to a multimode data row.

[0061] This ensures alignment of multimodal data 11 in both time and space, facilitating subsequent correlation analysis and fusion. For the audio data in each subgroup, speech recognition technology is used to extract the textual content, which is recorded as audio text. For video data, speech recognition is performed on the audio track to obtain audio recognition text; image recognition and scene understanding are performed on the video to generate image description text. These two components are then merged to form the complete video text, achieving semantic analysis of video content.

[0062] All text content is extracted from the subgroups, including original text data, audio text, and video text, and is comprehensively analyzed through natural language processing technology to generate a unified text summary.

[0063] Extract the file name, file address and metadata information of text data, audio data and video data respectively, including but not limited to: file name, storage path or access address, creation time, modification time, device number, resolution, sampling rate and other parameters.

[0064] Add a new row to the multimodal data table, assign it a unique row ID, and fill the extracted file name, file address, text summary, and various metadata information into the corresponding columns to form a complete multimodal data row.

[0065] For example, a construction company can group multimodal data 11 collected by multiple sensors deployed at multiple construction sites. This includes clock-in records, where employees clock in and out daily by scanning their faces, generating structured text-based clock-in records. Equipment management records record the usage time, operator, and operating status of various mechanical equipment, forming an equipment usage log. Cameras are deployed in key areas of the construction site to continuously record on-site conditions. Microphones are installed to capture on-site audio, such as worker conversations and unusual noises. All data for the day is grouped by source and then divided into sub-groups based on a preset cutoff period (for example, one hour).

[0066] For example, a subgroup includes the time period: 09:00 - 10:00, June 27, 2025; text data: tower crane clock-in records; text data: tower crane equipment usage records; video data: video surveillance (from the tower crane camera); and audio data: audio surveillance (from the tower crane microphone). These data sources are all tower crane-related and therefore grouped into the same subgroup. Table 1 shows an example row of multimode data in this embodiment.

[0067] Table 1 Example multi-mode data rows

[0068]

[0069] Through the above steps, the system automatically completes the collection, classification, identification, summary extraction, and structured storage of multimodal data, ultimately generating multimodal data rows in a unified format. This not only contains the basic information of the original data but also incorporates semantic descriptions, providing a high-quality data foundation for subsequent natural language-based intelligent retrieval, label matching, and semantic reasoning.

[0070] Step S2) Generate metadata description, content summary, preset scene tag set, and preset scene effect description for each multimodal data row based on the multimodal file description and multimodal file metadata.

[0071] The specific steps include:

[0072] generating a natural language task of generating a metadata description based on the multimodal file metadata and generating a content summary based on the multimodal file description;

[0073] Obtaining the metadata description and content summary based on the response of the pre-connected large language model 20 to the natural language task;

[0074] Reading scene descriptions of a plurality of preset scenes, and providing the scene descriptions to the large language model 20;

[0075] Submit a natural language task of generating a preset scene tag set and a preset scene effect description based on the scene description, metadata description and content summary to the large language model 20 to obtain the preset scene tag set and the preset scene effect description.

[0076] First, a natural language task is generated for the metadata in each multimodal data row (such as file name, address, creation time, etc.), converting these technical parameters into easily understandable descriptive text. Another natural language task is generated based on the multimodal file description (such as a summary of the video or audio content) to extract a concise and clear content summary. A large language model 20 processes these two natural language tasks, generating the corresponding metadata description and content summary, respectively.

[0077] The system reads descriptions of several preset typical application scenarios (such as "security monitoring", "equipment operation specification inspection", etc.) and inputs the scenario descriptions into the large language model 20 as context information.

[0078] The scenario descriptions of safety monitoring include: "Whether workers wear personal protective equipment correctly (such as helmets, gloves, reflective vests, etc.); whether there are unauthorized persons entering the construction area; whether there are dangerous behaviors or illegal operations (such as not using tools according to specifications, not wearing a safety rope when working at height, etc.); whether the on-site environment is stable, and whether there are safety hazards such as structural collapse and material sliding; detection of abnormal sounds or emergency call signals to respond to emergencies in a timely manner."

[0079] The scenario descriptions of equipment operation specification inspection include: "Operation qualification verification: confirm whether the operator performing a specific task has the corresponding qualification certificate and training experience; equipment status monitoring: regularly check the working status of mechanical equipment, including but not limited to key indicators such as temperature, pressure, vibration frequency, etc., to assess its health status and predict possible failures; Operation process compliance: supervise whether the operator strictly follows the predetermined operation manual to avoid equipment damage or safety accidents due to misoperation; Maintenance record review: verify whether the daily maintenance and repair records of the equipment are complete and accurate, and ensure that each maintenance work is properly performed; Abnormal behavior warning: use intelligent analysis algorithms to perform real-time analysis of the data stream generated during equipment operation. Once a deviation from the normal operating mode is detected (such as overload operation, unplanned shutdown, etc.), a warning message will be sent to the administrator immediately."

[0080] Submit a new natural language task to the large language model 20. This natural language task combines the previously generated metadata description, content summary, and current scene description to generate a label set suitable for a specific scene (such as "emergency", "normal operation") and a scene effect description (such as "no abnormal behavior was found", "potential safety hazards exist").

[0081] Taking multimodal data 11 from construction site monitoring as an example, the dataset read from the construction site includes: time clock text data: "Zhang San completed time clocking in at 9:00 on June 27, 2025." Equipment usage text data: "Tower crane X started operating at 9:15 on June 27, 2025." Video surveillance data: displays real-time footage of the tower crane area. Audio surveillance data: captures the voice communications between on-site workers regarding tower crane operations.

[0082] Generate a natural language task. For clock-in records, the generated natural language task is: "Generate a description based on the following metadata: User ID = ZhangSan, Timestamp = 2025-06-02T09:00:00, Location = Construction Site Entrance." For video surveillance, the generated natural language task is: "Generate a short content summary based on the following description: The video shows a tower crane carrying out material handling work."

[0083] Large language model 20 returns the metadata description: "Time 2025-06-02, punch-in record for construction site A"

[0084] Content summary: "Zhang San arrived at the construction site at the specified time and successfully signed in through the facial recognition system. Video footage shows that the tower crane was efficiently moving construction materials without any operational errors."

[0085] The scenario is described as follows: "On a construction site, it is crucial to ensure that all employees arrive on time and that machinery and equipment are functioning properly. In addition, it is necessary to closely monitor whether any violations of safety regulations occur."

[0086] Finally, the large language model 20 outputs a set of preset scenario labels: "Compliance Inspection," "Clock Clock Management," and "Construction Site Safety." The preset scenario descriptions include: "The construction site is operating normally, all employees have clocked in on time, tower crane operations are in accordance with standard procedures, and no violations have been observed."

[0087] Step S3) Generate guidance data 12 for each multi-mode data row, the guidance data 12 including a serial number, a starting storage address of the multi-mode data row, a metadata description, a content summary, a preset application scenario tag set, and a preset application scenario effect description.

[0088] Each Guide Data 12 entry will contain the following information:

[0089] Sequence number: used to identify the sequence number of the guide data 12 in the entire system.

[0090] Starting storage address: points to the specific location of the multi-mode data row in the storage system, facilitating quick location and access to the original data.

[0091] Metadata description: Provides basic descriptive information about the data, such as file name, creation time, source, etc.

[0092] Content summary: briefly summarize the main content or core points of the data to help users quickly understand the data overview.

[0093] Preset application scenario tag set: A set of tags assigned to data based on specific application scenarios (such as "security monitoring" and "equipment operation specification inspection") to facilitate classification and filtering.

[0094] Preset application scenario effect description: Based on the specific requirements of the application scenario, describe the performance or status of the data in the corresponding scenario, such as whether it complies with the specifications and whether there are any safety hazards.

[0095] By pre-generating guide data 12 for each multimodal data row, including information such as a sequence number, starting storage address, metadata description, content summary, preset application scenario tag set, and description of the preset application scenario effect, the system can quickly locate the data required by the user's query, effectively improving data retrieval speed. The metadata description and content summary in the guide data 12 provide a brief but comprehensive overview of the data, allowing users to understand its main content without having to access the entire original data. This is very useful for preliminary screening and determining which data is worth further review.

[0096] Step S4) The guide data 12 is stored in a pre-opened memory space and is persistently stored according to a preset period. When the multi-mode data 11 is updated, the multi-mode data row and the guide data 12 are updated.

[0097] A dedicated memory space needs to be allocated in advance for temporarily storing the guide data 12. This can significantly speed up the query process, as accessing data in memory is faster than reading it from disk.

[0098] Step S5) Receive a query requirement 30 expressed in natural language, and use the pre-connected large language model 20 to generate a tag set matching the query requirement 30.

[0099] Reading the scene descriptions of a plurality of preset scenes, and obtaining a total tag set according to the preset scene tag sets of all the guidance data 12;

[0100] Use the pre-connected large language model 20 to respectively calculate the semantic matching degree of the query requirement 30, the scene description of the preset scene, and each tag in the tag set;

[0101] According to the tags whose semantic matching degree is greater than a preset threshold, a tag set matching the query requirement 30 is obtained.

[0102] The query request is made in natural language, for example, such as "find all video records of equipment operation specification inspections in June 2025".

[0103] Based on the existing preset scene tag sets in all guidance data 12, the system constructs a set of all possible tags, called the tag set. The large language model 20 analyzes the user's query intent and compares it with each preset scene description to calculate the semantic similarity or match between the two. The relevance between the query requirement 30 and each tag in the tag set is further evaluated to obtain a specific semantic match score. To ensure the quality of the matching results, a minimum threshold for semantic matching is set (for example, 70%). Tags with a semantic match higher than the preset threshold are screened out to form a matching tag set for the current query requirement 30.

[0104] For example, a user may ask the following query: "Find all relevant information about construction site safety hazards in the past month."

[0105] Preset scenario descriptions: "Safety Monitoring", "Equipment Operation Specification Inspection".

[0106] Collection of tags: "Safety Hazards", "Illegal Operations", "Use of Personal Protective Equipment".

[0107] Semantic matching calculation: The two labels "security risk" and "illegal operation" are highly correlated with query requirement 30, with matching degrees of 90% and 80% respectively.

[0108] Filter to obtain the matching label set: Since the matching degree of both labels exceeds the preset threshold (assuming it is 75%), the final matching label set is ["security risk", "illegal operation"].

[0109] Through such a process, users' natural language queries can be effectively understood and relevant data resources can be accurately located, improving the efficiency and accuracy of data retrieval.

[0110] Step S6) Query the matching guide data 12 according to the tag set, obtain the matching multimodal data row 41 according to the matching guide data 12, and obtain the matching multimodal data 11 according to the multimodal data row.

[0111] The step of searching for matching guidance data 12 according to the tag set includes:

[0112] Calculating the semantic matching degree of each tag in the tag set with the content summary and the preset application scenario tag set in each of the guidance data 12;

[0113] According to the semantic matching degree between the tag and the content summary and the preset application scenario tag set, the matching guidance data 12 is obtained.

[0114] The system traverses all the guide data 12 and performs the following operations on each guide data 12:

[0115] a. Get the key fields in the guide data 12:

[0116] Content summary: Describes the core information of this multimodal data 11.

[0117] Preset application scenario label set: The data is labeled with application scenario-related labels (such as "security monitoring", "equipment operation specification inspection", "safety hazard", etc.).

[0118] b. Calculate semantic matching:

[0119] For each label in the label set (e.g., "safety hazard"), calculate its correlation with the current guidance data 12:

[0120] Semantic similarity with the content summary;

[0121] The degree of match with the preset application scenario label set.

[0122] c. Overall score:

[0123] The matching degrees of the two dimensions are integrated by using a weighted average or maximum value method to obtain the overall matching score between the label and the current guidance data 12.

[0124] A matching threshold (e.g., 0.7) is set. A piece of guide data 12 is considered a match only if its overall matching degree with at least one tag in the tag set exceeds this threshold. A set of matching guide data 12 is obtained, each of which matches the user query. Using the starting storage address in the guide data 12, the system can quickly locate the corresponding multimodal data row, thereby obtaining complete structured data information.

[0125] For example, a user enters a query requirement 30: "Find safety violations involving workers not wearing helmets in the past week." After processing, the system generates the following tag set: ["safety monitoring," "missing personal protective equipment," "violations"]. All guidance data 12 are traversed, and the semantic match between each guidance data 12 and the above tags is calculated. Take the following guidance data 12 as an example:

[0126] {

[0127] "Serial Number": "00127",

[0128] "Starting storage address": " / data / guide / 20250627_09",

[0129] "Content Summary": "The video shows a worker not wearing a helmet while working in the tower crane area."

[0130] "Preset application scenario label set": ["Security monitoring", "Violation"],

[0131] "Preset application scenario effect description": "Detecting that the helmet is not worn"

[0132] }

[0133] The matching degree between the guide data 12 and the labels is calculated as shown in Table 2. The maximum value is taken as the matching degree between the guide data 12 and the entire label set, which is 0.96. This is higher than the set threshold (0.7). Therefore, the guide data 12 is considered a match and added to the result set. The corresponding multimodal data row is found using the guide data 12, and the original video file and text record are further obtained.

[0134] Table 2 Matching degree between the guide data 12 and the label set

[0135]

[0136] Step S7) Generate a query result report 42 based on the metadata description, content summary, and preset application scenario effect description of the matched multi-mode data row 41.

[0137] The steps of generating the query result report 42 include:

[0138] Generate a basic information description of the matching multi-mode data row 41 according to the metadata description, wherein the basic information description includes a source description, a generation time description, a file type description, a file size description, and a file quantity description;

[0139] extracting a summary of the content summary as a summary description;

[0140] According to the preset application scenario effect description of the matched multi-mode data row 41, the effect description is superimposed to obtain the effect description;

[0141] Generate a query result report 42 containing the basic information description, summary description, and effect description.

[0142] Source description: Extracts the source information of the data from the metadata description, such as "surveillance camera at construction site A." A time description is generated, providing the specific time when the data was created or collected, such as "June 27, 2025, 9:00 AM." A file type description indicates the format of the data, such as "video file (MP4)," "audio file (WAV)," or "text record (TXT)." A file size description gives the file size, helping users understand the time and resources required to download or process it, such as "the video file size is 500MB." A file count description: If the matched data contains multiple files, the number of files is indicated, such as "a total of 3 related video files were found."

[0143] Extract the content summary from the matching multimodal data row 41 as a brief but comprehensive description, allowing users to understand the core information without viewing the entire data. For example, "The video shows a worker not wearing a helmet while working in the tower crane area."

[0144] In the effect description obtained by superimposing the preset application scenario effect description, the preset application scenario effect description in the matching multi-mode data row 41 is used to provide a more in-depth effect analysis. For example, in the "safety monitoring" scenario, the effect description may be "Detected that the person was not wearing a helmet, posing a safety hazard." This provides a description of the effect of the potential hazard for reference.

[0145] All the above information is integrated together to form a complete query result report42.

[0146] On the other hand, this manual provides an intelligent data query system based on big data, please refer to the attached Figure 4 ,include:

[0147] The reading module 100 reads the multimodal data 11 and creates a multimodal data row according to the source and acquisition time. The multimodal data row includes a row ID, a multimodal file name, a multimodal file address, a multimodal file description, and multimodal file metadata.

[0148] A first generating module 200 generates a metadata description, a content summary, a preset scene tag set, and a preset scene effect description for each multimodal data row based on the multimodal file description and the multimodal file metadata;

[0149] A second generating module 300 generates guide data 12 for each multi-mode data row, wherein the guide data 12 includes a sequence number, a starting storage address of the multi-mode data row, a metadata description, a content summary, a preset application scenario tag set, and a preset application scenario effect description;

[0150] The guidance module 400 stores the guidance data 12 in a pre-opened memory space and stores it persistently according to a preset period. When the multi-mode data 11 is updated, it triggers the update of the multi-mode data row and the guidance data 12.

[0151] The third generation module 500 receives a query requirement 30 expressed in natural language and generates a tag set matching the query requirement 30 using the pre-connected large language model 20;

[0152] A query module 600 queries the matching guide data 12 according to the tag set, obtains the matching multimodal data row 41 according to the matching guide data 12, and obtains the matching multimodal data 11 according to the multimodal data row;

[0153] The result module 700 generates a query result report 42 based on the metadata description, content summary, and preset application scenario effect description of the matched multi-mode data row 41.

[0154] See also Figure 5 The figure shows a schematic diagram of the structure of an electronic device provided by an embodiment of this specification.

[0155] like Figure 5As shown, the electronic device 1100 may include: at least one processor 1101, at least one network interface 1104, a user interface 1103, a memory 1105, and at least one communication bus 1102. The communication bus 1102 may be used to implement communication between the aforementioned components. The user interface 1103 may include buttons, and optionally may also include a standard wired interface or a wireless interface. The network interface 1104 may include, but is not limited to, a Bluetooth module, an NFC module, a Wi-Fi module, etc. The processor 1101 may include one or more processing cores. The processor 1101 utilizes various interfaces and circuits to connect the various components within the electronic device 1100. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1105, and accessing data stored in the memory 1105, the processor 1101 performs various functions of the routing device and processes data. Optionally, the processor 1101 may be implemented in hardware using at least one of a DSP, an FPGA, and a PLA. The processor 1101 may integrate one or a combination of a CPU, a GPU, and a modem. Among them, the CPU mainly processes the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content that needs to be displayed on the display; and the modem is used to handle wireless communication.

[0156] It is understandable that the above-mentioned modem may not be integrated into the processor 1101, but may be implemented separately through a chip.

[0157] Memory 1105 may include either RAM or ROM. Optionally, memory 1105 may include non-transitory computer-readable media. Memory 1105 may be used to store instructions, programs, codes, code sets, or instruction sets. Memory 1105 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, sound playback function, image playback function, etc.), instructions for implementing the aforementioned method embodiments, etc.; the data storage area may store data related to the aforementioned method embodiments, etc. Memory 1105 may also optionally be at least one storage device located remotely from the aforementioned processor 1101. Memory 1105, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and application programs. Processor 1101 may be configured to invoke the application programs stored in memory 1105 and execute the methods described in the aforementioned embodiments.

[0158] The embodiments of this specification also provide a computer-readable storage medium having instructions stored therein that, when executed on a computer or processor, cause the computer or processor to perform the steps of the aforementioned embodiments. If the components of the aforementioned electronic device are implemented as software functional units and sold or used as independent products, they may be stored in the computer-readable storage medium.

[0159] The embodiments of this specification also provide a computer program product, including a computer program, which implements multiple steps in the above embodiments when executed by a processor.

[0160] In the absence of conflict, the technical features in this embodiment and implementation scheme can be combined arbitrarily.

[0161] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product comprises multiple computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center that integrates multiple available media. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).

[0162] When implemented via hardware or firmware, the aforementioned method flow is programmed into the hardware circuit to obtain the corresponding hardware circuit structure and realize the corresponding function. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit, whose logical function is determined by the user's device programming. Designers can "integrate" a digital system on a PLD through self-programming, eliminating the need for chip manufacturers to design and manufacture dedicated integrated circuit chips. Moreover, today, instead of manually manufacturing integrated circuit chips, this programming is often performed using "logic compiler" software. This is similar to the software compiler used in program development. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There are not just one HDL, but many. Those skilled in the art will also understand that simply by programming the method flow in one of the aforementioned hardware description languages ​​and programming it into the integrated circuit, a hardware circuit that implements the logical method flow can be easily obtained.

[0163] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Without departing from the design spirit of this specification, various modifications and improvements made to the technical solutions of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims of this specification.

Claims

1. An intelligent data query method based on big data, characterized in that: Including steps: Read multimodal data and create a multimodal data row based on the source and acquisition time. The multimodal data row includes a row ID, a multimodal file name, a multimodal file address, a multimodal file description, and multimodal file metadata. Generating a metadata description, a content summary, a preset scene tag set, and a preset scene effect description for each multimodal data row according to the multimodal file description and the multimodal file metadata; Generate guidance data for each multi-mode data row, the guidance data including a sequence number, a starting storage address of the multi-mode data row, a metadata description, a content summary, a preset application scenario tag set, and a preset application scenario effect description; The guidance data is stored in a pre-opened memory space and is persistently stored according to a preset period. When the multi-mode data is updated, the multi-mode data row and the guidance data are triggered to be updated; Receive a query requirement expressed in natural language, and use a pre-connected large language model to generate a tag set that matches the query requirement; querying matching guide data according to the tag set, obtaining matching multimodal data rows according to the matching guide data, and obtaining matching multimodal data according to the multimodal data rows; Generate a query result report based on the metadata description, content summary, and preset application scenario effect description of the matched multi-mode data row; The step of generating a metadata description, a content summary, a preset scene tag set, and a preset scene effect description for each multimodal data row according to the multimodal file description and the multimodal file metadata includes: generating a natural language task of generating a metadata description based on the multimodal file metadata and generating a content summary based on the multimodal file description; Obtaining the metadata description and content summary based on a response of a pre-connected large language model to the natural language task; Reading scene descriptions of a plurality of preset scenes, and providing the scene descriptions to the large language model; Submitting a natural language task of generating a preset scene tag set and a preset scene effect description based on the scene description, metadata description, and content summary to the large language model to obtain the preset scene tag set and the preset scene effect description; The application scenario effect is described as the performance or status of the data in the corresponding scenario.

2. The intelligent data query method based on big data according to claim 1, characterized in that: The multimodal data includes text data, audio data and video data, The steps to create multi-mode data rows based on source and acquisition time include: Establishing a multimodal data table, wherein the columns of the multimodal data table include an ID column, a text table column, a text address column, a text description column, a text metadata column, an audio table column, an audio address column, an audio description column, an audio metadata column, a video table column, a video address column, a video description column, and a video metadata column; The multi-mode data is grouped according to the source, and the multi-mode data is truncated according to the generation time according to the preset truncation period. The multi-mode data of the same group in the same truncation period is regarded as a sub-group, and one sub-group corresponds to one multi-mode data row; Identify the text in the sub-grouped audio data and record it as audio text; Identify the sound recognition text of the audio track in the video data, identify the picture of the video, and obtain the picture description text, wherein the sound recognition text and the picture description text constitute the video text; Extracting text summaries of text data, audio text, and video text in the subgroups; Extracting the file names, file addresses and metadata of the text data, audio data and video data respectively; Add a new row to the multi-mode data table, assign a unique row ID to the new row and fill it in the ID column; Fill the file names, file addresses, text summaries and metadata corresponding to the text data, audio data and video data into the corresponding columns.

3. The intelligent data query method based on big data according to claim 1 or 2, characterized in that: The steps of receiving a query requirement expressed in natural language and generating a tag set matching the query requirement using a pre-connected large language model include: Read the scene descriptions of several preset scenes, and obtain the total tag set according to the preset scene tag set of all guidance data; Use the pre-connected large language model to calculate the semantic matching degree of the query requirement, the scene description of the preset scene, and each tag in the total tag set; A tag set matching the query requirement is obtained based on the tags whose semantic matching degree is greater than a preset threshold.

4. The intelligent data query method based on big data according to claim 1 or 2, characterized in that: The step of searching for matching guidance data according to the tag set includes: Calculating the semantic matching degree of each tag in the tag set with the content summary and the preset application scenario tag set in each of the guidance data respectively; According to the semantic matching degree between the tag and the content summary and the preset application scenario tag set, matching guidance data is obtained.

5. The intelligent data query method based on big data according to claim 1 or 2, characterized in that: The steps of generating a query result report based on the metadata description, content summary, and preset application scenario effect description of the matched multi-mode data row include: Generate a basic information description of the matching multi-mode data row according to the metadata description, wherein the basic information description includes a source description, a generation time description, a file type description, a file size description, and a file quantity description; extracting a summary of the content summary as a summary description; According to the preset application scenario effect descriptions of the matched multi-mode data rows, superimpose to obtain an effect description; Generate a query result report containing the basic information description, summary description, and effect description.

6. Intelligent data query system based on big data, characterized by: include: A reading module reads multimodal data and creates a multimodal data row according to the source and acquisition time. The multimodal data row includes a row ID, a multimodal file name, a multimodal file address, a multimodal file description, and multimodal file metadata. A first generation module generates a metadata description, a content summary, a preset scene tag set, and a preset scene effect description for each multimodal data row based on the multimodal file description and the multimodal file metadata; a second generating module, generating guide data for each multi-mode data row, wherein the guide data includes a sequence number, a starting storage address of the multi-mode data row, a metadata description, a content summary, a preset application scenario tag set, and a preset application scenario effect description; A guidance module stores the guidance data in a pre-opened memory space and stores it persistently according to a preset period. When the multi-mode data is updated, the multi-mode data row and the guidance data are triggered to be updated. A third generation module receives a query requirement expressed in natural language and uses a pre-connected large language model to generate a tag set that matches the query requirement; a query module, querying matching guide data according to the tag set, obtaining matching multimodal data rows according to the matching guide data, and obtaining matching multimodal data according to the multimodal data rows; The result module generates a query result report based on the metadata description, content summary, and preset application scenario effect description of the matched multi-modal data rows; The step of generating a metadata description, a content summary, a preset scene tag set, and a preset scene effect description for each multimodal data row according to the multimodal file description and the multimodal file metadata includes: generating a natural language task of generating a metadata description based on the multimodal file metadata and generating a content summary based on the multimodal file description; Obtaining the metadata description and content summary based on a response of a pre-connected large language model to the natural language task; Reading scene descriptions of a plurality of preset scenes, and providing the scene descriptions to the large language model; Submitting a natural language task of generating a preset scene tag set and a preset scene effect description based on the scene description, metadata description, and content summary to the large language model to obtain the preset scene tag set and the preset scene effect description; The application scenario effect is described as the performance or status of the data in the corresponding scenario.

7. An electronic device, characterized in that: including a processor and a memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • General security risk monitoring method and system based on large model capability

    CN119964083A

  • Multi-modal data generation and fine adjustment method for domain image

    CN120046117A