A Multimodal Data Retrieval Method Based on Semantic Association
By using a customized deep neural network model to intelligently identify videos recorded in law enforcement scenarios, the problem of quickly retrieving primary crime scene videos from legal documents and case databases by law enforcement agencies has been solved, achieving efficient and accurate intelligent identification and reducing manual retrieval time.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING FOREST POLICE COLLEGE
- Filing Date
- 2025-07-17
- Publication Date
- 2026-05-26
Smart Images

Figure CN120849655B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent integrated circuits, and more specifically relates to the field of big data processing and artificial intelligence technology, particularly to a multimodal data retrieval method based on semantic association. Background Technology
[0002] Big data processing with artificial intelligence refers to the use of advanced technologies and algorithms to deeply analyze and mine large-scale data to extract valuable information and insights. It combines big data technology, artificial intelligence (AI), machine learning (ML), and data mining methods to automate the analysis of complex datasets, discover potential value and correlations, and achieve automated data processing and analysis to support decision-making and optimize business processes. Compared to traditional manual analysis, big data processing with AI features automation, in-depth mining, real-time processing, and visualization. Intelligent big data analytics is widely used in various fields, including financial services, healthcare, retail, and marketing, helping businesses make more accurate decisions and enhance their competitiveness.
[0003] For example, Chinese invention patent publication CN117033724A proposes a multimodal data retrieval method based on semantic association, relating to the field of multimodal data retrieval technology. The method includes the following steps: collecting modal data information and retrieval evaluation index information during the operation of a multimodal data retrieval system based on semantic association; performing comprehensive analysis to generate an accuracy evaluation index; establishing a dataset merging mechanism; performing comprehensive analysis on the accuracy evaluation indices within the dataset; generating operational status signals; and issuing different prompts accordingly. This invention evaluates the accuracy of semantic association modeling in a multimodal data retrieval system. When accuracy decreases, the system promptly detects this and prompts relevant maintenance personnel to take corresponding maintenance and optimization measures. This ensures the accuracy of semantic association modeling, guarantees that the model effectively captures the semantic associations between data, effectively prevents the relevance of the system's returned retrieval results to the user's query, and effectively prevents providing users with misleading retrieval results.
[0004] For example, Chinese invention patent publication CN118051653A proposes a multimodal data retrieval method, system, and medium based on semantic association. The method includes: acquiring data to be retrieved in different modalities and performing semantic recognition; generating semantic feature data for different modalities and classifying them to generate multimodal semantic feature data clusters of different categories; performing semantic recognition on user retrieval data and processing it to obtain extended search keyword data; matching and recognizing the extended search keyword data with the multimodal semantic feature data clusters to obtain retrieval result data; processing user historical search log data to obtain a user search preference model; performing cross-modal semantic fusion on the retrieval result data; and processing it according to the user search preference model to obtain optimized semantic fusion result data and sending it to the user; thereby achieving the goal of rapid retrieval of multimodal data and accurately adapting to user search needs.
[0005] However, the aforementioned technical solutions only address user queries and retrieval based on semantic association, lacking targeted solutions for key application scenarios. For example, law enforcement agencies need to analyze multiple data sources such as semantics, images, and videos, and utilize deep learning and pattern recognition technologies to improve their understanding and response capabilities to complex scenarios. This would enable rapid retrieval of legal documents and case databases, obtaining multimodal data retrieval results based on semantic association. For instance, multimodal data retrieval based on semantic association would be required to sequentially perform intelligent identification on recorded videos from legal documents and case databases that match a specific law enforcement scenario to determine whether they belong to the primary crime scene, until a recorded video matching a specific law enforcement scenario belonging to the primary crime scene is retrieved. This would provide crucial information for law enforcement agencies in enforcing the law and making judgments regarding that specific law enforcement scenario. Clearly, the aforementioned technical solutions cannot accomplish this task. Summary of the Invention
[0006] When law enforcement agencies investigate each law enforcement scene, they will take photos of that scene from various angles. At the same time, multiple law enforcement scenes will be photographed during the law enforcement process, resulting in a huge amount of video recordings of various scenes. These videos are stored in the law enforcement agencies' legal documents and case databases. Specifically, due to differences in shooting angles and content, some video recordings may be considered secondary crime scenes, while others may be primary crime scenes but the shooting angles may not meet the requirements for people and evidence at the primary crime scene.
[0007] In this way, law enforcement agencies, during subsequent enforcement investigations and evidence collection, need to meticulously examine a large number of scene-recorded videos to retrieve key information, such as those belonging to the primary crime scene. This primary crime scene video not only meets the requirements for people and evidence at the primary crime scene in terms of shooting angle, but more importantly, it is not a secondary crime scene or other fabricated scene. This provides crucial information for law enforcement agencies to conduct enforcement and judgment for each scenario, eliminating the need for manual, one-by-one examination of numerous scene-recorded videos. This invention provides a solution for the intelligent and rapid retrieval of scene-recorded videos belonging to the primary crime scene.
[0008] Specifically, this invention provides a multimodal data retrieval method based on semantic association. Addressing the need for law enforcement agencies to quickly retrieve video recordings of each law enforcement scenario that belong to the primary crime scene from legal documents and case databases, this method employs a customized artificial intelligence model. Based on semantic, image, and video data derived from the video recordings, it intelligently determines whether each video recording belongs to the primary crime scene. If it does not, it retrieves the next video recording of that law enforcement scenario and continues intelligently until a video recording of the primary crime scene within that law enforcement scenario is found. This completes the multimodal data retrieval based on semantic association for each law enforcement scenario, providing crucial information for law enforcement agencies in their enforcement and judgment for each scenario.
[0009] According to the present invention, a multimodal data retrieval method based on semantic association is provided, the method comprising:
[0010] Retrieve scene recording videos that match actual law enforcement scenarios from legal documents and case databases belonging to law enforcement agencies. Each scene recording video that matches a law enforcement scenario has a set number of multi-frame scene recording images.
[0011] The set of existing strings corresponding to each frame of the scene recording video matched with the actual law enforcement scene is used as the semantic association information of the scene recording video matched with the actual law enforcement scene.
[0012] The recording frame rate, video resolution, average number of targets, and average brightness value of the video recorded in the actual law enforcement scenario are used as video association information.
[0013] Each frame of the scene recording video matched with the actual law enforcement scene is used as the image association information of the scene recording video.
[0014] The deep neural network, after multiple training sessions, intelligently identifies whether a video recorded in a scene matched with an actual law enforcement scenario belongs to the primary crime scene based on the semantic association information, video association information, and image association information of a set number of such videos and the actual law enforcement scenario.
[0015] If the scene recording video matched by the intelligent identification of the actual law enforcement scenario does not belong to the primary crime scene, the system continues to search for the next scene recording video that matches the actual law enforcement scenario from the legal documents and case database belonging to the law enforcement department; otherwise, the search is stopped.
[0016] Among them, the deep neural network after multiple training sessions is the intelligent identification model, and the number of training sessions of the deep neural network is positively correlated with the set number.
[0017] Compared with the prior art, the present invention has at least the following four key inventive points:
[0018] Invention Point 1: Retrieves video recordings of actual law enforcement scenarios from legal documents and case databases belonging to law enforcement agencies. Each video recording of a law enforcement scenario has a set number of multi-frame scene recording images. The semantic association information of the video recording is based on the set of existence strings corresponding to each frame of the scene recording images. The video association information is based on the recording frame rate, video resolution, average number of existing targets, and average brightness value of the video recording. The image association information is based on the set of customized visual data corresponding to each frame of the scene recording images. This completes the multimodal data retrieval of video recordings of actual law enforcement scenarios based on semantic association, providing sufficient and comprehensive basic data for subsequent intelligent identification of whether the video recordings of actual law enforcement scenarios belong to the primary crime scene.
[0019] Invention Point Two: To determine whether a video recording of a scene matched with an actual law enforcement scenario belongs to the primary crime scene, a customized intelligent identification model is designed. The intelligent identification model is a deep neural network that has undergone multiple training iterations, and the number of training iterations of the deep neural network is positively correlated with a set number, thereby ensuring the reliability and effectiveness of the intelligent identification.
[0020] Invention Point 3: When the scene recording video matched by the intelligent identification of the actual law enforcement scenario does not belong to the first crime scene, the system continues to search for the next scene recording video that matches the actual law enforcement scenario from the legal documents and case database belonging to the law enforcement department. Otherwise, the search stops. This allows for targeted retrieval of related videos of the first crime scene for each law enforcement scenario using a polling retrieval method based on semantic association. This provides accurate information support for each law enforcement scenario quickly and saves a lot of manual retrieval and retrieval time.
[0021] Invention Point 4: In each training iteration of the deep neural network, the known scene identifier indicating whether a video recording of a certain scene matched by a certain law enforcement scenario belongs to the first crime scene is used as the output content of the deep neural network, and the set number of semantic association information, video association information and image association information of the video recording of the certain scene matched by the certain law enforcement scenario are used as the input content of the deep neural network to complete the training, thereby ensuring the training effect of each iteration.
[0022] Invention Point 5: The deep neural network used includes a single output layer, a single input layer, and multiple hidden layers, wherein the multiple hidden layers are located between the single output layer and the single input layer, and the number of hidden layers in the deep neural network is positively correlated with a set number, i.e., the number of frames of multi-frame scene recording images. Attached Figure Description
[0023] The embodiments of the present invention will now be described with reference to the accompanying drawings, wherein:
[0024] Figure 1 This is a schematic diagram illustrating the working scenario of the semantic association-based multimodal data retrieval method according to the present invention.
[0025] Figure 2 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 1 of the present invention.
[0026] Figure 3 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 2 of the present invention.
[0027] Figure 4 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 3 of the present invention.
[0028] Figure 5 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 4 of the present invention.
[0029] Figure 6The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 5 of the present invention. Detailed Implementation
[0030] like Figure 1 The diagram shows a working scenario of a multimodal data retrieval method based on semantic association according to the present invention. The big data processing artificial intelligence technology of the present invention belongs to the field of intelligent integrated circuits.
[0031] When law enforcement agencies investigate each law enforcement scene, they will take photos of that scene from various angles. At the same time, they will take photos of multiple law enforcement scenes during the law enforcement process, resulting in a huge amount of video recordings of various scenes, which are stored in the law enforcement agencies' legal documents and case databases.
[0032] In this way, law enforcement agencies, during subsequent enforcement investigations and evidence collection, need to meticulously examine a large number of scene-recorded videos to retrieve key information, such as those belonging to the primary crime scene. This primary crime scene video not only meets the requirements for people and evidence at the primary crime scene in terms of shooting angle, but more importantly, it is not a secondary crime scene or other fabricated scene. This provides crucial information for law enforcement agencies to conduct enforcement and judgment for each scenario, eliminating the need for manual, one-by-one examination of numerous scene-recorded videos. This invention provides a solution for the intelligent and rapid retrieval of scene-recorded videos belonging to the primary crime scene.
[0033] Therefore, the specific technical process of the present invention is as follows:
[0034] Technical Process A: Retrieve scene recording videos that match actual law enforcement scenarios from legal documents and case databases belonging to law enforcement agencies. Each scene recording video matching a law enforcement scenario has a set number of multi-frame scene recording images.
[0035] Specifically, such as Figure 1 The video recordings of the scenes shown may be considered a second crime scene due to different shooting angles and content, or they may be from the first crime scene but the shooting angle does not meet the requirements for people and evidence of the first crime scene.
[0036] Technical Process B: Utilizing each actual law enforcement scenario-matched video recording retrieved through Technical Process A, analyze its corresponding semantic association information, video association information, and image association information to serve as multimodal data retrieval information based on semantic association, such as... Figure 1 As shown;
[0037] Specifically, the set of existing strings corresponding to each frame of the scene recording video matched with the actual law enforcement scene serves as the semantic association information of the scene recording video matched with the actual law enforcement scene.
[0038] Specifically, the recording frame rate, video resolution, average number of targets, and average brightness value of the scene-recorded videos matched with actual law enforcement scenarios are used as video association information.
[0039] Specifically, each frame of the scene recording video matched with the actual law enforcement scenario is used as the image association information of the scene recording video.
[0040] In this way, multimodal data retrieval of scene recording videos matched with actual law enforcement scenarios is completed based on semantic association, providing sufficient and comprehensive basic data for subsequent intelligent identification of whether scene recording videos matched with actual law enforcement scenarios belong to the primary crime scene;
[0041] Technical Process C: To determine whether a video recording of a scene matched with an actual law enforcement scenario constitutes the primary crime scene, a customized intelligent identification model is designed.
[0042] For example, the structural customization of the intelligent identification model is mainly reflected in the following aspects:
[0043] First: The intelligent identification model is a deep neural network that has undergone multiple training iterations;
[0044] Second: The number of training iterations of the deep neural network is positively correlated with the set number;
[0045] For example, if the number is set to 100, the corresponding number of training iterations for the deep neural network is 600; if the number is set to 200, the corresponding number of training iterations for the deep neural network is 700; if the number is set to 300, the corresponding number of training iterations for the deep neural network is 800; if the number is set to 400, the corresponding number of training iterations for the deep neural network is 900, and so on.
[0046] Third: In each training session of the deep neural network, the known scene identifier indicating whether a certain law enforcement scene matched with a certain scene recording video belongs to the first crime scene is used as the output content of the deep neural network, and the set number of semantic association information, video association information and image association information of the certain law enforcement scene matched with the certain scene recording video are used as the input content of the deep neural network to complete the training, thereby ensuring the training effect of each training session.
[0047] Fourth: The deep neural network includes a single output layer, a single input layer, and multiple hidden layers, and in the deep neural network, the multiple hidden layers are located between the single output layer and the single input layer, including: the number of hidden layers in the deep neural network is positively correlated with the set number;
[0048] For example, if the set number is 100 frames, the number of hidden layers of the deep neural network positively associated with the set number is 3; if the set number is 200 frames, the number of hidden layers of the deep neural network positively associated with the set number is 4; if the set number is 300 frames, the number of hidden layers of the deep neural network positively associated with the set number is 5; if the set number is 400 frames, the number of hidden layers of the deep neural network positively associated with the set number is 6, and so on.
[0049] Thus, through the aforementioned customized structural design of the intelligent identification model, the reliability and effectiveness of intelligent identification are ensured.
[0050] Technical Process D: The intelligent identification model, using a customized structure designed according to Technical Process C, retrieves multimodal data of scene recording videos matching actual law enforcement scenarios based on semantic association in Technical Process B. This data enables intelligent identification of whether the scene recording videos matching actual law enforcement scenarios belong to the primary crime scene. Figure 1 As shown, a scene identifier is obtained to indicate whether the scene recording video matching the actual law enforcement scenario belongs to the primary crime scene;
[0051] Technical Process E: Based on the intelligent identification results of Technical Process D, determine whether it is necessary to continue searching the legal documents and case databases belonging to law enforcement agencies for the next scene recording video that matches the actual law enforcement scenario, thereby completing the intelligent retrieval of the scene recording video of the first crime scene required for law enforcement.
[0052] For example, if the scene recording video matched by the intelligent identification of the actual law enforcement scenario does not belong to the first crime scene, the system continues to search for the next scene recording video that matches the actual law enforcement scenario from the legal documents and case database belonging to the law enforcement department; otherwise, the search stops.
[0053] Therefore, through the coordinated operation of the above-mentioned technical processes, the present invention can complete the targeted retrieval of related videos of the first crime scene in various law enforcement scenarios using a polling retrieval method based on semantic association, thereby providing accurate information support for various law enforcement scenarios and saving a lot of manual retrieval and retrieval time.
[0054] The key points of this invention are: targeted screening of semantic association information, video association information and image association information of scene recording videos; multimodal data retrieval of scene recording videos matched with actual law enforcement scenarios based on semantic association; multiple customized structure design of intelligent identification model; and targeted design of polling retrieval method based on semantic association.
[0055] The semantic association-based multimodal data retrieval method of the present invention will be described in detail below by way of embodiments.
[0056] Example 1
[0057] Figure 2 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 1 of the present invention.
[0058] like Figure 2 As shown, the multimodal data retrieval method based on semantic association includes the following specific steps:
[0059] Step 21: Retrieve scene recording videos that match actual law enforcement scenarios from the legal documents and case database belonging to law enforcement agencies. Each scene recording video matching a law enforcement scenario has a set number of multi-frame scene recording images.
[0060] Specifically, due to differences in shooting angle and content, scene recording videos may be considered a second crime scene, or they may be at the first crime scene but the shooting angle does not meet the requirements for people and evidence at the first crime scene.
[0061] For example, in a specific case, there are scene recording videos for various scenarios, such as the violence scene, the indirect scene, and the fabricated scene. Each scenario, as a law enforcement scenario, involves multiple scene recording videos taken from different angles by law enforcement investigators during their on-site investigation. These massive amounts of scene recording videos are stored in the law enforcement department's legal documents and case database. During subsequent law enforcement and evidence collection, it is necessary to quickly retrieve the scene recording videos belonging to the primary crime scene from the law enforcement department's legal documents and case database. Generally, scene recording videos belonging to the primary crime scene come from the violence scene. However, there are also requirements for the shooting angle of scene recording videos belonging to the primary crime scene to ensure that the people and evidence at the primary crime scene are present in the video footage. If the massive amount of scene recording videos in the law enforcement department's legal documents and case database were manually searched one by one to obtain the required scene recording videos belonging to the primary crime scene, the workload would be enormous. This invention is designed to solve this technical problem.
[0062] Step 22: The set of existing strings corresponding to each frame of the scene recording video matched with the actual law enforcement scene is used as the semantic association information of the scene recording video matched with the actual law enforcement scene.
[0063] Specifically, the set of existence strings corresponding to each frame of scene recording image indicates which semantic information can be successfully identified in the frame of scene recording image;
[0064] More specifically, the identification of semantic information present in each frame of the scene recording image is usually performed using OCR semantic recognition mode;
[0065] Step 23: Use the recording frame rate, video resolution, average number of targets, and average brightness value of the scene recording video matched with the actual law enforcement scene as video association information;
[0066] For example, the recording frame rate, video resolution, average number of targets, and average brightness value of the scene recording video matched with the actual law enforcement scene are used as video association information. The recording frame rate of the scene recording video matched with the actual law enforcement scene can be selected as 20 frames per second, 25 frames per second, 30 frames per second, or 50 frames per second.
[0067] Step 24: Use the customized visual data corresponding to each frame of the scene recording video matched with the actual law enforcement scene as the image association information of the scene recording video matched with the actual law enforcement scene.
[0068] Specifically, the image association information of the scene recording video matched with the actual law enforcement scene is usually the visualization information of the scene recording video matched with the actual law enforcement scene, such as brightness values, etc.
[0069] Step 25: Using a deep neural network that has undergone multiple training iterations, the system intelligently identifies whether a video recording of a scene matched with an actual law enforcement scenario belongs to the primary crime scene based on the semantic association information, video association information, and image association information of the video recordings, which are matched with a set number of actual law enforcement scenarios.
[0070] Step 26: If the scene recording video matched by the intelligent identification of the actual law enforcement scenario does not belong to the primary crime scene, continue to search for the next scene recording video that matches the actual law enforcement scenario from the legal documents and case database belonging to the law enforcement department; otherwise, stop searching.
[0071] In this way, the present invention can complete the targeted retrieval of related videos of the first crime scene in various law enforcement scenarios using a polling retrieval method based on semantic association, thereby providing accurate information support for various law enforcement scenarios quickly and saving a lot of manual retrieval and retrieval time.
[0072] Among them, the deep neural network after multiple training sessions is the intelligent identification model, and the number of training sessions of the deep neural network is positively correlated with the set number.
[0073] For example, a deep neural network that has undergone multiple training iterations is an intelligent identification model, and the number of training iterations of the deep neural network is positively correlated with a set number, including: when the set number is 100, the corresponding number of training iterations of the deep neural network is 600; when the set number is 200, the corresponding number of training iterations of the deep neural network is 700; when the set number is 300, the corresponding number of training iterations of the deep neural network is 800; when the set number is 400, the corresponding number of training iterations of the deep neural network is 900, and so on.
[0074] Among them, the set of existing strings corresponding to each frame of the scene recording video matched with the actual law enforcement scene is the set of strings obtained by concatenating the first and last strings in the scene recording image according to their order of appearance in the scene recording image.
[0075] Among them, the set of existing strings corresponding to each frame of the scene recording video matched by the actual law enforcement scene is the set of strings that exist in the scene recording image in that frame, which is obtained by connecting the first and last characters of each string in the scene recording image in the order of their appearance in the scene recording image in that frame. This includes: when the number of characters in the set of existing strings corresponding to the scene recording image in that frame is greater than a set number threshold, the last character of the set of existing strings corresponding to the scene recording image in that frame is pruned to obtain a set of existing strings with a number of characters equal to the set number threshold.
[0076] Among them, the set of existing strings corresponding to each frame of the scene recording video matched by the actual law enforcement scene is the set of strings that exist in the scene recording image in that frame, which is obtained by connecting the first and last strings in the high and low order of their appearance in the scene recording image in that frame. It also includes: when the number of characters in the set of existing strings corresponding to the scene recording image in that frame is less than a set number threshold, zero-padding is performed on the end of the set of existing strings corresponding to the scene recording image in that frame to obtain a set of existing strings with a number of characters equal to the set number threshold.
[0077] In each training iteration of the deep neural network, the known scene identifier indicating whether a video recording of a certain scene matched by a certain law enforcement scenario belongs to the first crime scene is used as the output of the deep neural network, and the set number of semantic association information, video association information, and image association information of the video recording of the certain scene matched by the certain law enforcement scenario are used as the input of the deep neural network to complete the training.
[0078] Specifically, the MATLAB toolbox can be used to complete the testing and simulation of the training process of this training operation by using the known scene identifiers that indicate whether a certain law enforcement scene matches a certain scene recording video belongs to the first crime scene as the output content of the deep neural network, and using the set number of semantic association information, video association information and image association information of the certain scene recording video that matches the certain law enforcement scene as the input content of the deep neural network.
[0079] The process of using a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scenario belongs to the first crime scene involves: simultaneously inputting the semantic association information, video association information, and image association information of a set number of videos recording the scene matched with an actual law enforcement scenario into the deep neural network that has undergone multiple training iterations, and executing the deep neural network that has undergone multiple training iterations to obtain a scene identifier output by the deep neural network that has undergone multiple training iterations, indicating whether the video recording the scene matched with an actual law enforcement scenario belongs to the first crime scene.
[0080] Specifically, the number of targets in each frame of the scene recording video matched with the actual law enforcement scene is obtained, the maximum and minimum values of each target number are removed to obtain the remaining multiple target numbers, and the arithmetic mean of the multiple target data is taken as the average number of targets in the scene recording video matched with the actual law enforcement scene.
[0081] Specifically, the brightness values of each frame of the scene recording video matched with the actual law enforcement scene are obtained, the maximum and minimum values of each set of overall image brightness values are removed to obtain the remaining multiple sets of overall image brightness values, and the arithmetic mean of the multiple sets of overall image brightness values is taken as the average brightness value of the scene recording video matched with the actual law enforcement scene.
[0082] The process of obtaining the number of existing targets corresponding to each frame of the scene recording video matched with the actual law enforcement scene, removing the maximum and minimum values from each number of existing targets to obtain the remaining number of existing targets, and taking the arithmetic mean of the multiple number of existing targets as the average number of existing targets in the scene recording video matched with the actual law enforcement scene includes: each existing target in each frame of the scene recording image is each image block remaining after the background area of the frame of the scene recording image is stripped.
[0083] Specifically, the background region of each frame of scene recording image can be identified based on the imaging features of the background region. Then, the background region is stripped from each frame of scene recording image, and the remaining image blocks are used as the existing targets in that frame of scene recording image.
[0084] In addition, the customized visual data corresponding to each frame of the scene recording video matched with the actual law enforcement scene is the depth value, brightness value, horizontal coordinate value and vertical coordinate value of each pixel of the scene recording image.
[0085] Example 2
[0086] Figure 3 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 2 of the present invention.
[0087] like Figure 3 As shown, with Figure 2 Unlike the previous embodiment, before retrieving scene recording videos matching actual law enforcement scenarios from legal documents and case databases belonging to law enforcement agencies, and before each scene recording video matching a law enforcement scenario has a set number of multi-frame scene recording images, i.e. before step S21, the method further includes:
[0088] Step S27: Use Alibaba PolarDB database to create a legal document and case library belonging to law enforcement agencies;
[0089] For example, different scene recordings are stored at different physical addresses in legal documents and case repositories belonging to law enforcement agencies.
[0090] Example 3
[0091] Figure 4 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 3 of the present invention.
[0092] like Figure 4 As shown, with Figure 2 Unlike the previous embodiment, before using a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scene belongs to the first crime scene based on a set number of semantic association information, video association information, and image association information, i.e. before step S25, the method further includes:
[0093] Step S28: Use a model storage device to store the deep neural network model after multiple training iterations;
[0094] Among them, the use of model storage devices to store the model of the deep neural network after multiple trainings includes: storing the various model parameters of the deep neural network after multiple trainings, i.e., the intelligent identification model, to complete the model storage of the deep neural network after multiple trainings.
[0095] Specifically, setting the number to 100 corresponds to 600 training iterations for the deep neural network; setting the number to 200 corresponds to 700 training iterations for the deep neural network; setting the number to 300 corresponds to 800 training iterations for the deep neural network; setting the number to 400 corresponds to 900 training iterations for the deep neural network, and so on.
[0096] Example 4
[0097] Figure 5 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 4 of the present invention.
[0098] like Figure 5 As shown, with Figure 2 Unlike the previous embodiment, after using a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scene belongs to the first crime scene based on a set number of semantic association information, video association information, and image association information, i.e. after step S25, the method further includes:
[0099] Step S29: If the scene recording video matched by the intelligent identification of the actual law enforcement scenario belongs to the first crime scene, the key video of the scene recording video matched by the actual law enforcement scenario is marked as the first crime scene and stored in the legal documents and case library belonging to the law enforcement department; if the scene recording video matched by the intelligent identification of the actual law enforcement scenario does not belong to the first crime scene, the general video of the scene recording video matched by the actual law enforcement scenario is marked as the non-first crime scene and stored in the legal documents and case library belonging to the law enforcement department.
[0100] In this way, the present invention can complete the targeted retrieval of related videos of the first crime scene in various law enforcement scenarios using a polling retrieval method based on semantic association, thereby providing accurate information support for various law enforcement scenarios quickly and saving a lot of manual retrieval and retrieval time.
[0101] Example 5
[0102] Figure 6 The following is a flowchart illustrating the steps of a multimodal data retrieval method based on semantic association according to Embodiment 5 of the present invention.
[0103] like Figure 6 As shown, with Figure 2 Unlike the previous embodiment, after using a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scene belongs to the first crime scene based on a set number of semantic association information, video association information, and image association information, i.e. after step S25, the method further includes:
[0104] Step S30: If the scene video matched by the intelligent identification of the actual law enforcement scene belongs to the first crime scene, the scene video matched by the actual law enforcement scene is compressed and wirelessly transmitted to the supervision server of the remote law enforcement supervision department; if the scene video matched by the intelligent identification of the actual law enforcement scene does not belong to the first crime scene, the scene video matched by the actual law enforcement scene is not compressed and wirelessly transmitted.
[0105] For example, if the intelligent identification of a scene recording video matched with an actual law enforcement scenario belongs to the first crime scene, the scene recording video matched with the actual law enforcement scenario is compressed and wirelessly transmitted to the supervision server of the remote law enforcement supervision department; and if the intelligent identification of a scene recording video matched with an actual law enforcement scenario does not belong to the first crime scene, the scene recording video matched with the actual law enforcement scenario is not compressed and wirelessly transmitted, including: the wireless transmission is based on a time-division duplex communication link or a frequency-division duplex communication link.
[0106] Next, the various method embodiments of the present invention will be described in detail.
[0107] In the semantic association-based multimodal data retrieval method according to various method embodiments of the present invention:
[0108] The semantic association information, video association information, and image association information of the set number of scene recording videos matched with the actual law enforcement scenario are synchronously input into the deep neural network after multiple training sessions, and the deep neural network after multiple training sessions is executed to obtain the scene identifier of the deep neural network after multiple training sessions, which indicates whether the scene recording video matched with the actual law enforcement scenario belongs to the first crime scene. This includes: performing binary numerical conversion on the set number of scene recording videos matched with the actual law enforcement scenario, the semantic association information, video association information, and image association information, and then synchronously inputting them into the deep neural network after multiple training sessions.
[0109] For example, in the semantic association information, video association information and image association information of the video recordings of the set number of actual law enforcement scenarios, the semantic association information may itself be the binary numerical representation of ASCII code, so there is no need to perform binary numerical conversion. However, for the video association information and image association information that are not binary numerical representations, binary numerical conversion needs to be performed separately to complete the normalization processing of the model input content.
[0110] The process involves simultaneously inputting the semantic association information, video association information, and image association information of the scene recording videos matched with the actual law enforcement scenario into a deep neural network that has undergone multiple training iterations. The deep neural network is then executed to obtain a scene identifier output by the deep neural network that indicates whether the scene recording video matched with the actual law enforcement scenario belongs to the first crime scene. This also includes a scene identifier in binary numerical form indicating whether the scene recording video matched with the actual law enforcement scenario belongs to the first crime scene.
[0111] And in the semantic association-based multimodal data retrieval method according to various method embodiments of the present invention:
[0112] The semantic association information, video association information, and image association information of a set number of scene-recorded videos matching actual law enforcement scenarios are converted into binary values and then synchronously input into a deep neural network that has undergone multiple training iterations. This includes: using a numerical conversion device to perform binary numerical conversion on the semantic association information, video association information, and image association information of a set number of scene-recorded videos matching actual law enforcement scenarios.
[0113] For example, using a numerical conversion device to perform binary numerical conversion on the semantic association information, video association information, and image association information of a set number of scene-recorded videos matching actual law enforcement scenarios includes: the numerical conversion device can be a programmable logic device designed in VHDL language;
[0114] The process of converting the semantic association information, video association information, and image association information of the set number of scene-recorded videos matched with the actual law enforcement scenario into binary values and then synchronously inputting them into the deep neural network after multiple training sessions also includes: using a synchronous control device to synchronously input the semantic association information, video association information, and image association information of the set number of scene-recorded videos matched with the actual law enforcement scenario into the deep neural network after multiple training sessions, after the binary value conversion is performed.
[0115] For example, using a synchronization control device to synchronously input a set number of binary numerical conversions, semantic association information of scene-recorded videos matched with actual law enforcement scenarios, video association information, and image association information into a deep neural network that has undergone multiple training iterations includes: the synchronization control device can be a SOC chip or an ASIC chip;
[0116] Among them, the deep neural network after multiple training sessions is an intelligent identification model and the number of training sessions of the deep neural network is positively correlated with the set number includes: using an information mapping formula to represent the information mapping relationship that the deep neural network after multiple training sessions is an intelligent identification model and the number of training sessions of the deep neural network is positively correlated with the set number.
[0117] The information mapping formula used to represent the information mapping relationship between the deep neural network after multiple training sessions as an intelligent identification model and the number of training sessions of the deep neural network being positively correlated with a set number includes: in the information mapping formula, the set number is the input information of the information mapping formula;
[0118] Furthermore, the information mapping formula used to represent the deep neural network after multiple training sessions as an intelligent identification model, and the information mapping relationship in which the number of training sessions of the deep neural network is positively correlated with a set number, also includes: in the information mapping formula, the number of training sessions of the deep neural network positively correlated with the set number is the output information of the information mapping formula.
[0119] Furthermore, the present invention may also reference the following technical contents to further demonstrate the outstanding substantial progress of the present invention:
[0120] The deep neural network includes a single output layer, a single input layer, and multiple hidden layers, wherein the multiple hidden layers are located between the single output layer and the single input layer in the deep neural network;
[0121] The deep neural network includes a single output layer, a single input layer, and multiple hidden layers. In the deep neural network, the multiple hidden layers are located between the single output layer and the single input layer, meaning that the number of hidden layers in the deep neural network is positively correlated with the predetermined number.
[0122] For example, the number of hidden layers in the deep neural network being positively correlated with the set number includes: when the set number is 100 frames, the number of hidden layers in the deep neural network being positively correlated with the set number is 3; when the set number is 200 frames, the number of hidden layers in the deep neural network being positively correlated with the set number is 4; when the set number is 300 frames, the number of hidden layers in the deep neural network being positively correlated with the set number is 5; when the set number is 400 frames, the number of hidden layers in the deep neural network being positively correlated with the set number is 6, and so on.
[0123] The deep neural network includes a single output layer, a single input layer, and multiple hidden layers. In the deep neural network, the multiple hidden layers are located between the single output layer and the single input layer. The system further includes a numerical conversion formula that represents a positive correlation between the number of hidden layers in the deep neural network and the set number.
[0124] The deep neural network includes a single output layer, a single input layer, and multiple hidden layers. Furthermore, in the deep neural network, the multiple hidden layers are located between the single output layer and the single input layer. Additionally, in the numerical conversion formula, the set number is the input value of the numerical conversion formula, and the number of hidden layers positively correlated with the set number is the output value of the numerical conversion formula.
[0125] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal to execute the methods described in the various embodiments of the present invention.
[0127] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. A multimodal data retrieval method based on semantic association, characterized in that, The method includes: Retrieve scene recording videos that match actual law enforcement scenarios from legal documents and case databases belonging to law enforcement agencies. Each scene recording video that matches a law enforcement scenario has a set number of multi-frame scene recording images. The set of existing strings corresponding to each frame of the scene recording video matched with the actual law enforcement scene is used as the semantic association information of the scene recording video matched with the actual law enforcement scene. The recording frame rate, video resolution, average number of targets, and average brightness of the video recorded in the actual law enforcement scenario are used as video association information. Each frame of the scene recording video matched with the actual law enforcement scenario is used as the image association information of the scene recording video. The deep neural network, after multiple training sessions, intelligently identifies whether a video recorded in a scene matched with an actual law enforcement scenario belongs to the primary crime scene based on the semantic association information, video association information, and image association information of a set number of such videos and the actual law enforcement scenario. If the scene recording video matched by the intelligent identification of the actual law enforcement scenario does not belong to the primary crime scene, the system continues to search for the next scene recording video that matches the actual law enforcement scenario from the legal documents and case database belonging to the law enforcement department; otherwise, the search is stopped. Among them, the deep neural network after multiple training sessions is the intelligent identification model, and the number of training sessions of the deep neural network is positively correlated with the set number. Among them, the set of existing strings corresponding to each frame of the scene recording video matched with the actual law enforcement scene is the set of strings obtained by concatenating the first and last strings in the scene recording image according to their order of appearance in the scene recording image. In each training iteration of the deep neural network, the known scene identifier indicating whether a video recording of a certain scene matched with a certain law enforcement scenario belongs to the first crime scene is used as the output of the deep neural network. The set number of semantic association information, video association information, and image association information of the video recording of the certain scene matched with the certain law enforcement scenario are used as the input of the deep neural network to complete the training.
2. The multimodal data retrieval method based on semantic association as described in claim 1, characterized in that: The set of existing strings corresponding to each frame of the scene recording video matched in the actual law enforcement scenario is the set of strings obtained by concatenating the strings existing in that frame of scene recording image in the order of their appearance in that frame of scene recording image. This includes: when the number of characters in the set of existing strings corresponding to that frame of scene recording image is greater than a set number threshold, the last character of the set of existing strings corresponding to that frame of scene recording image is pruned to obtain a set of existing strings with a number of characters equal to the set number threshold. The set of existing strings corresponding to each frame of the scene recording video matched in the actual law enforcement scenario is a set of strings obtained by concatenating the strings existing in that frame of scene recording image in the order of their appearance in that frame of scene recording image. It also includes: when the number of characters in the set of existing strings corresponding to that frame of scene recording image is less than a set number threshold, zero-padding is performed on the end of the set of existing strings corresponding to that frame of scene recording image to obtain a set of existing strings with a number of characters equal to the set number threshold.
3. The multimodal data retrieval method based on semantic association as described in claim 2, characterized in that: The method employs a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scenario belongs to the primary crime scene, based on a set number of such videos and semantic association information, video association information, and image association information. This involves simultaneously inputting a set number of such videos into the deep neural network that has undergone multiple training iterations, and executing the deep neural network to obtain a scene identifier output by the deep neural network that indicates whether the video recording of the scene matched with the actual law enforcement scenario belongs to the primary crime scene.
4. The multimodal data retrieval method based on semantic association as described in claim 2, characterized in that: Obtain the number of targets in each frame of the scene recording video matched with the actual law enforcement scene. Remove the maximum and minimum values from each number of targets to obtain the remaining number of targets. Use the arithmetic mean of the remaining number of targets as the average number of targets in the scene recording video matched with the actual law enforcement scene. Specifically, the brightness values of each frame of the scene recording video matched with the actual law enforcement scene are obtained, the maximum and minimum values of each set of overall image brightness values are removed to obtain the remaining multiple sets of overall image brightness values, and the arithmetic mean of the multiple sets of overall image brightness values is taken as the average brightness value of the scene recording video matched with the actual law enforcement scene. The process of obtaining the number of existing targets corresponding to each frame of the scene recording video matched with the actual law enforcement scene, removing the maximum and minimum values from each number of existing targets to obtain the remaining number of existing targets, and taking the arithmetic mean of the multiple number of existing targets as the average number of existing targets in the scene recording video matched with the actual law enforcement scene includes: each existing target in each frame of the scene recording image is each image block remaining after the background area of the frame of the scene recording image is stripped. Among them, the customized visual data corresponding to each frame of the scene recording video matched with the actual law enforcement scenario is the depth value, brightness value, horizontal coordinate value and vertical coordinate value of each pixel of the scene recording image.
5. The multimodal data retrieval method based on semantic association as described in claim 3, characterized in that, Before retrieving scene recording videos matching actual law enforcement scenarios from legal documents and case databases belonging to law enforcement agencies, wherein each scene recording video matching a law enforcement scenario has a predetermined number of multi-frame scene recording images, the method further includes: The Alibaba PolarDB database is used to create a legal document and case library belonging to law enforcement agencies.
6. The multimodal data retrieval method based on semantic association as described in claim 3, characterized in that, Before using a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scenario belongs to the primary crime scene based on semantic association information, video association information, and image association information of a set number of such videos, the method further includes: A model storage device is used to store the model of a deep neural network after multiple training iterations; The method of storing the model of a deep neural network after multiple training sessions using a model storage device includes storing the various model parameters of the deep neural network after multiple training sessions, i.e., the intelligent identification model.
7. The multimodal data retrieval method based on semantic association as described in claim 3, characterized in that, After employing a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scenario belongs to the primary crime scene based on semantic association information, video association information, and image association information from a set number of such videos and matching actual law enforcement scenarios, the method further includes: If the intelligent identification system matches the recorded video of the actual law enforcement scenario and it belongs to the primary crime scene, the key video of the recorded video of the actual law enforcement scenario matching the primary crime scene will be stored in the legal documents and case library belonging to the law enforcement department. If the intelligent identification system matches the recorded video of the actual law enforcement scenario and it does not belong to the primary crime scene, the general video of the recorded video of the actual law enforcement scenario matching the primary crime scene will be stored in the legal documents and case library belonging to the law enforcement department.
8. The multimodal data retrieval method based on semantic association as described in claim 3, characterized in that, After employing a deep neural network that has undergone multiple training iterations to intelligently identify whether a video recording of a scene matched with an actual law enforcement scenario belongs to the primary crime scene based on semantic association information, video association information, and image association information from a set number of such videos and matching actual law enforcement scenarios, the method further includes: If the video recording of the scene matched by the intelligent identification system in the actual law enforcement scenario belongs to the primary crime scene, the video recording of the scene matched by the actual law enforcement scenario will be compressed and wirelessly transmitted to the supervision server of the remote law enforcement supervision department. If the video recording of the scene matched by the intelligent identification system in the actual law enforcement scenario does not belong to the primary crime scene, the video recording of the scene matched by the actual law enforcement scenario will not be compressed or wirelessly transmitted.
9. The multimodal data retrieval method based on semantic association as described in any one of claims 4-8, characterized in that: The semantic association information, video association information, and image association information of the set number of scene recording videos matched with the actual law enforcement scenario are synchronously input into the deep neural network after multiple training sessions, and the deep neural network after multiple training sessions is executed to obtain the scene identifier of the deep neural network after multiple training sessions, which indicates whether the scene recording video matched with the actual law enforcement scenario belongs to the first crime scene. This includes: performing binary numerical conversion on the set number of scene recording videos matched with the actual law enforcement scenario, the semantic association information, video association information, and image association information, and then synchronously inputting them into the deep neural network after multiple training sessions. The process involves simultaneously inputting the semantic association information, video association information, and image association information of the scene recording videos matched with the actual law enforcement scenario into a deep neural network that has undergone multiple training iterations. The deep neural network is then executed to obtain a scene identifier output by the deep neural network that indicates whether the scene recording video matched with the actual law enforcement scenario belongs to the first crime scene. This also includes a scene identifier in binary numerical form indicating whether the scene recording video matched with the actual law enforcement scenario belongs to the first crime scene.
10. The multimodal data retrieval method based on semantic association as described in any one of claims 4-8, characterized in that: The semantic association information, video association information, and image association information of a set number of scene-recorded videos matching actual law enforcement scenarios are converted into binary values and then synchronously input into a deep neural network that has undergone multiple training iterations. This includes: using a numerical conversion device to perform binary numerical conversion on the semantic association information, video association information, and image association information of a set number of scene-recorded videos matching actual law enforcement scenarios. The process of converting the semantic association information, video association information, and image association information of the set number of scene-recorded videos matched with the actual law enforcement scenario into binary values and then synchronously inputting them into the deep neural network after multiple training sessions also includes: using a synchronous control device to synchronously input the semantic association information, video association information, and image association information of the set number of scene-recorded videos matched with the actual law enforcement scenario into the deep neural network after multiple training sessions, after the binary value conversion is performed. Among them, the deep neural network after multiple training sessions is an intelligent identification model and the number of training sessions of the deep neural network is positively correlated with the set number includes: using an information mapping formula to represent the information mapping relationship that the deep neural network after multiple training sessions is an intelligent identification model and the number of training sessions of the deep neural network is positively correlated with the set number. The information mapping formula used to represent the information mapping relationship between the deep neural network after multiple training sessions as an intelligent identification model and the number of training sessions of the deep neural network being positively correlated with a set number includes: in the information mapping formula, the set number is the input information of the information mapping formula; The information mapping formula, which expresses the information mapping relationship that the deep neural network after multiple training sessions is an intelligent identification model and that the number of training sessions of the deep neural network is positively correlated with a set number, further includes: in the information mapping formula, the number of training sessions of the deep neural network that is positively correlated with the set number is the output information of the information mapping formula.