A multimodal target data intelligent processing system

Through the multimodal target data intelligent processing system, combined with the feature analysis of video, image and text data, the problem of low storage and analysis efficiency of target data in the existing technology is solved, and more efficient and accurate data correlation analysis and storage management are achieved.

CN119597981BActive Publication Date: 2025-08-26NANJING XIAOWANLI INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510125250.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-08-26
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

The prior art is inefficient and inaccurate in the classification storage and analysis of target data, especially when comprehensively processing video data, image data and text data, it is not possible to achieve effective correlation analysis.

Method used

A multimodal target data intelligent processing system is adopted, including a target data acquisition module, a video data extraction module, an image information analysis module, a personnel information analysis module and a data storage management module. Through the analysis of video frame change characteristics, image local features, personnel correlation parameters, etc., the feature correlation analysis and storage management of video, image and text data is realized.

Benefits of technology

It improves the accuracy of correlation analysis between data, enhances the analysis efficiency of target data and the accuracy of storage processing, and ensures the efficiency of data retrieval and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119597981B_ABST
    Figure CN119597981B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing technology, and in particular to a multimodal target data intelligent processing system, comprising: a target data acquisition module for acquiring target scene information and associated personnel information; a video data extraction module for extracting the target area of ​​the video, analyzing the video frame change characteristics, and extracting the scene video image; an image information analysis module for analyzing the local features and the overall features of the image, and analyzing the degree of prominence of the local features; a personnel information analysis module for storing associated personnel information, counting the number of text matches and the number of matching personnel, and analyzing personnel association parameters; and a data storage management module for analyzing feature correlation and storing the target scene information and associated personnel information based on the feature correlation. The present invention realizes intelligent storage and processing of target data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multi-modal target data intelligent processing system. Background Art

[0002] Target data processing relies heavily on the development of big data technology. With the explosive growth of data volumes, proper data classification, storage, and processing have become increasingly important. By leveraging big data technology to collect, store, process, and analyze target data, we can rationally segment massive amounts of data and improve data processing efficiency and quality.

[0003] Chinese Patent Publication No. CN117763174A discloses a multimodal retrieval method, device, and storage medium, including: determining image information and text information as retrieval input information; encoding the text information using a pre-set first encoding module to determine a first text feature belonging to a text feature space corresponding to the text information; migrating the first text feature using a pre-set first migration module to determine a corresponding second text feature belonging to an image feature space, and merging the first text feature and the second text feature to generate a third text feature; encoding the image information to generate an image feature corresponding to the image information; merging the third text feature with the image feature to generate a combined image-text feature; and performing a search based on the combined image-text feature to obtain a search result corresponding to the image information and the text information. This invention implements a classified analysis of content features in image data, but does not implement the classified storage of target data by integrating video data, image data, and text data. This leads to low efficiency in target data analysis and inaccurate storage and processing of target data. Summary of the Invention

[0004] An object of the present invention is to provide a multimodal target data intelligent processing system to solve at least one of the problems existing in the prior art.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A multimodal target data intelligent processing system, comprising:

[0007] Target data collection module, used to collect target site information and related personnel information;

[0008] A video data extraction module is used to extract the target area of ​​the video according to the target scene information, and is also used to analyze the change characteristics of the video frame according to the target scene information, and extract the scene video image according to the video frame change characteristics and the video target area;

[0009] An image information analysis module is used to analyze the local features and overall features of the image based on the target scene information, and is also used to analyze the degree of local feature significance based on the local features of the image;

[0010] The personnel information analysis module is used to store the associated personnel information and match the associated personnel information with the preset extracted keywords to count the number of text matches and the number of matching personnel, and also to analyze the personnel association parameters based on the number of text matches and the number of matching personnel;

[0011] The data storage management module is used to analyze the feature correlation based on the local features of the image, the overall features of the image, the on-site video image and the personnel association parameters, and store the target scene information and the associated personnel information according to the feature correlation.

[0012] Furthermore, the video data extraction module is provided with a change feature analysis unit, which is used to extract an image frame every K frames from the on-site monitoring video as a video frame image, and number the video frame images according to the extraction order to obtain a video frame number, where K represents a video extraction parameter;

[0013] The change feature analysis unit analyzes the video frame change feature according to the video frame image to obtain the video frame change feature Q(k), where k represents the video frame number.

[0014] Furthermore, the video data extraction module is also provided with a video image extraction unit, which is used to extract the on-site video image according to the video frame change feature Q(k). If there is a target area in the on-site monitoring video and Q(k)≤q, the video image extraction unit currently analyzes the video frame image corresponding to the video frame change feature as the on-site video image; otherwise, the video image extraction unit does not extract the on-site video image; wherein q represents the change feature threshold.

[0015] Furthermore, the image information analysis module is provided with a division region analysis unit, which is used to analyze the local features of the image according to the image split regions to obtain the local features of the image W(v), where v represents the split number;

[0016] The divided region analysis unit analyzes the local feature obviousness according to the local feature W(v) of the image, and the local feature obviousness includes the local feature not obvious and the local feature obvious.

[0017] Furthermore, the image information analysis module is also provided with a picture feature analysis unit, which is used to analyze the overall image features based on the on-site grayscale image to obtain the overall image features A(i), where i represents the on-site grayscale image number.

[0018] Furthermore, the personnel information analysis module matches the target text with the preset extraction keywords, and counts the number of times the text extraction keywords appear in the target text as the text matching number N1(j). If N1(j)>0, the personnel information analysis module sets the current analysis-related personnel information as the matching personnel; if N1(j)=0, the personnel information analysis module sets the current analysis-related personnel information as the unmatched personnel; the personnel information analysis module counts the number of associated personnel information of the matching personnel as the matching personnel number NR2, and analyzes the personnel association parameter according to the text matching number N1(j) and the matching personnel number NR2 to obtain the personnel association parameter R(j), where j represents the personnel number.

[0019] Furthermore, the data storage management module is provided with a feature correlation analysis unit, which is used to analyze the feature correlation degree according to the overall image feature A(i), the on-site video image and the personnel correlation parameter R(j) to obtain the feature correlation degree F(m).

[0020] Furthermore, the data storage management module is also provided with a local feature matching unit, which is used to take the split number corresponding to the image split area with obvious local features as the obvious feature number set U1, and analyze the local feature correlation based on the obvious feature number set U1, the local feature W(v) of the image and the overall feature A(i) of the image. The local feature correlation includes weak local feature correlation and strong local feature correlation, and when the local feature correlation is strong, the feature correlation degree analysis process is processed, and the feature correlation degree after processing is F1(m).

[0021] Furthermore, the data storage management module is further provided with a storage matching analysis unit, which is used to count the number of stored target scene information and associated personnel information that satisfy |F(m) / F(z)-1|≤α1 as the storage matching number NF, and extract the stored target scene information that satisfies |A(i) / A(z)-1|≤α1 and has the same photo type as the matching scene information, and the storage matching analysis unit counts the number of matching scene information of the first storage type as the first matching number NA1, and counts the number of matching scene information of the second storage type as the second matching number NA2, wherein F(z) represents the feature correlation degree corresponding to the stored target scene information and the associated personnel information, z represents the storage information number, A(z) represents the overall image feature corresponding to the stored target scene information, and α1 represents the first storage matching threshold;

[0022] The storage matching analysis unit analyzes the storage correlation according to the storage matching number NF, the first matching number NA1 and the second matching number NA2, wherein the storage correlation includes strong storage correlation and weak storage correlation, and processes the analysis process of the local feature matching parameters when the storage correlation is strong, and the processed local feature matching parameters are B1.

[0023] Furthermore, the data storage management module is further provided with a data storage management unit, which is used to analyze the storage type of the target site information and the associated personnel information according to the feature correlation degree, and the storage type includes the first category and the second category.

[0024] The beneficial effects of the present invention are as follows: through the target data acquisition module to collect target site information and related personnel information, and other modules to analyze the collected data, the storage of the collected data is managed by integrating video data, image data and text data, and the analysis of the correlation between each data feature and the data is realized, and the accuracy of the correlation analysis between the data is improved, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1 Schematic diagram of the structure of the multimodal target data intelligent processing system of this embodiment.

[0027] Figure 2 Schematic diagram of the structure of the video data extraction module in this embodiment.

[0028] Figure 3 Schematic diagram of the structure of the image information analysis module of this embodiment.

[0029] Figure 4 This is a structural diagram of the data storage management module in this embodiment. DETAILED DESCRIPTION

[0030] In order to more clearly illustrate the present invention, the present invention is further described below in conjunction with preferred embodiments and accompanying drawings. Similar components in the accompanying drawings are represented by the same reference numerals. It should be understood by those skilled in the art that the following detailed description is illustrative rather than restrictive and should not be used to limit the scope of protection of the present invention.

[0031] It should be noted that, although the terms "first," "second," and "third" may be used to describe the embodiments of the present application, the description should not be limited to these terms. These terms are merely used to distinguish the descriptions. For example, without departing from the scope of the embodiments of the present application, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first."

[0032] See also Figure 1 As shown, it is a multimodal target data intelligent processing system of this embodiment, including:

[0033] The target data acquisition module is used to collect target scene information and related personnel information. The target scene information includes the scene location, scene photo information and scene surveillance video. The scene photo information includes photo type and scene grayscale image. The photo type is data about the content of the scene photo. The scene grayscale image is the grayscale image of the scene photo. The scene surveillance video is the surveillance video data of the target scene. The related personnel information includes personnel identity information and target text. The personnel identity information includes but is not limited to name, gender, age, ID number and other data. The target scene information and related personnel information are collected by user interactive input.

[0034] Specifically, the target scene information in this embodiment is the scene data of a target location analyzed by the system, and the associated personnel information is the personnel information that may be associated with the target scene.

[0035] Specifically, this embodiment is applied to a cloud server storing target data, and by analyzing the correlation between target data and data characteristics, it realizes classified storage of target data, thereby ensuring the efficiency of subsequent retrieval and analysis of target data.

[0036] It should be noted that the target information (including but not limited to target device information, target personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the target or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant national and regional laws, regulations and standards.

[0037] Please continue reading Figure 1 As shown, the multimodal target data intelligent processing system also includes:

[0038] The video data extraction module is used to extract the video target area according to the target scene information, and is also used to analyze the video frame change characteristics according to the target scene information, and extract the scene video image according to the video frame change characteristics and the video target area. The video data extraction module is connected to the target data acquisition module.

[0039] Specifically, in this embodiment, when analyzing image data, the pixel point in the lower left corner of the image is used as the coordinate origin, and the two sides adjacent to the coordinate origin are used as the x-axis and y-axis respectively. The x-axis increases from left to right, and the y-axis increases from bottom to top. The unit length is 1 pixel. A plane rectangular coordinate system is established, and the coordinate points are used to represent the position of each pixel in the image.

[0040] See also Figure 2 As shown, the video data extraction module includes:

[0041] The target area extraction unit is used to extract the video target area according to the on-site monitoring video, and the video target area includes the human body area and the target area.

[0042] Specifically, the target region extraction unit in this embodiment uses a target detection model to extract the video target region, and has extracted the video target region in the on-site monitoring video.

[0043] Specifically, the target detection model described in this embodiment is a target detection model based on YoloV8. The target detection model is trained using a human detection dataset and a vehicle detection dataset to achieve the extraction of human areas and vehicle areas in on-site surveillance videos. This embodiment does not specifically limit the setting of the detection dataset, such as using the VOC dataset and the COCO dataset.

[0044] Specifically, in this embodiment, the target area extraction unit analyzes the on-site surveillance video to extract the target area in the on-site surveillance video, thereby realizing the extraction of people and vehicles in the on-site surveillance video, ensuring the accuracy of target extraction in the surveillance video data, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0045] Please continue reading Figure 2 As shown, the video data extraction module also includes:

[0046] The change feature analysis unit is used to analyze the change features of the video frame according to the on-site monitoring video, and the change feature analysis unit is connected to the target area extraction unit.

[0047] Specifically, the change feature analysis unit described in this embodiment extracts an image frame every K frames from the on-site surveillance video as a video frame image, and numbers the video frame images in the order in which they were extracted to obtain a video frame number, where K represents a video extraction parameter, and 20≤K≤30. It will be appreciated that this embodiment does not impose specific limitations on the value of the video extraction parameter, and those skilled in the art may freely set it as long as it satisfies the extraction of video frame images. The optimal value of the video extraction parameter is: K = 27.

[0048] Specifically, the change feature analysis unit described in this embodiment analyzes the video frame change feature based on the video frame image, sets the video frame change feature to Q(k), and sets Q(k)=σ1(k) / [σ1(k-1)+1], where σ1(k) represents the image feature of the video frame image, and k represents the video frame number.

[0049] Specifically, the image feature of the video frame image in this embodiment is the standard deviation of the grayscale value of the pixel in the video frame image, which is calculated as follows: , L k (x, y) represents the grayscale value of each pixel in the video frame image, NXY represents the number of pixels in the video frame image, Avg(L k (x, y) represents the average grayscale value of the pixels in the video frame image. It is understood that the setting of the image features of the video frame image is not specifically limited in this embodiment, and those skilled in the art may freely set the image features. For example, the image features of the video frame image may be set to image features calculated using image feature extraction algorithms such as SIFT and HOG.

[0050] Specifically, in this embodiment, the change feature analysis unit extracts the video frame image to extract part of the data used for analysis in the surveillance video, thereby analyzing the video frame change features, and using the video frame change features to represent the differences in image features between different video frames, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0051] Please continue reading Figure 2 As shown, the video data extraction module also includes:

[0052] The video image extraction unit is used to extract the on-site video image according to the video frame change characteristics and the video target area. The video image extraction unit is connected to the change characteristic analysis unit.

[0053] Specifically, the video image extraction unit in this embodiment extracts live video images based on video frame change characteristics. If a target area exists in the live surveillance video and Q(k) ≤ q, the video image extraction unit analyzes the video frame image corresponding to the video frame change characteristics as the live video image. Otherwise, the video image extraction unit does not extract the live video image. Wherein, q represents the change characteristic threshold, and 0.5 ≤ q ≤ 0.8. It is understood that the value of the change characteristic threshold is not specifically limited in this embodiment, and those skilled in the art can freely set it as long as it satisfies the extraction of live video images. The optimal value of the change characteristic threshold is: q = 0.6.

[0054] Specifically, in this embodiment, the video image extraction unit analyzes the video change characteristics and video target areas to extract images in the on-site monitoring video where the detection target exists and the extracted video frame images have large differences in change as on-site video images, thereby ensuring the diversity of the extracted on-site video images, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0055] Please continue reading Figure 1 As shown, the multimodal target data intelligent processing system also includes:

[0056] The image information analysis module is used to analyze the local features and the overall features of the image according to the target scene information. The image information analysis module is connected to the video data extraction module.

[0057] See also Figure 3 As shown, the image information analysis module includes:

[0058] The scene image splitting unit is used to split the scene grayscale image to obtain image split areas.

[0059] Specifically, the scene image segmentation unit in this embodiment creates a rectangular area of ​​pixel size u×u, and aligns the lower left corner of the rectangular area with the lower left corner pixel of the scene grayscale image. The corresponding pixel in the rectangular area is extracted as an image segmentation area, and the rectangular area is translated to the right by one pixel after each extraction. If the rectangular area is translated to the right to the right boundary of the scene grayscale image, the rectangular area is translated upward by one pixel. If the rectangular area is translated to the left to the left boundary of the scene grayscale image, the rectangular area is translated upward by one pixel until the rectangular area is translated to the left to the upper left boundary of the scene grayscale image or to the right to the upper right boundary of the scene grayscale image. In this way, (NX-u)×(NY-u) image segmentation areas are extracted and numbered in the order of extraction to obtain a segmentation number, where u represents a segmentation parameter, 4≤u≤10, NX represents the number of pixels in the x-axis direction of the scene grayscale image, and NY represents the number of pixels in the y-axis direction of the scene grayscale image. The optimal value of the segmentation parameter is u=5.

[0060] Specifically, in this embodiment, the scene grayscale image is split by the scene image splitting unit to obtain multiple image splitting areas, thereby increasing the diversity of the system's analysis of the scene grayscale image and improving the accuracy of the analysis of each area in the scene grayscale image, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0061] Please continue reading Figure 3 As shown, the image information analysis module also includes:

[0062] The divided area analysis unit is used to analyze the local features of the image according to the image split areas, and to analyze the degree of obviousness of the local features according to the local features of the image. The divided area analysis unit is connected to the on-site image splitting unit.

[0063] Specifically, the divided region analysis unit described in this embodiment analyzes the local features of the image according to the image split region, sets the local features of the image to W(v), and sets W(v)=σ2(v), where σ2(v) represents the image features of the image split region, and v represents the split number.

[0064] Specifically, the image features of the image split regions described in this embodiment are analyzed in the same manner as the image features of the video frame images, and will not be further elaborated in this embodiment.

[0065] Specifically, the region division analysis unit in this embodiment analyzes the degree of local feature prominence based on the local features of the image. If W(v) / Avg(L i (x, y))≤w, the divided region analysis unit determines that the local features of the current analysis image split region are not obvious; otherwise, the divided region analysis unit determines that the local features of the current analysis image split region are obvious; wherein Avg(L i (x, y) represents the average grayscale value of the pixels in the scene grayscale image, and w represents the visibility threshold, with 0.05≤w≤0.15. It is understood that the value of the visibility threshold is not specifically limited in this embodiment, and those skilled in the art may freely set it, as long as it satisfies the analysis of the local feature visibility. The optimal value of the visibility threshold is: w = 0.1.

[0066] Specifically, in this embodiment, the image split area is analyzed by the divided area analysis unit to analyze the local features of the image, and the image features of each image split area are represented by the local features of the image, so as to analyze the degree of local feature significance and realize the extraction of the image split area with obvious local features, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0067] Please continue reading Figure 3 As shown, the image information analysis module also includes:

[0068] The picture feature analysis unit is used to analyze the overall features of the image based on the on-site photo information. The picture feature analysis is connected to the divided area analysis unit.

[0069] Specifically, the image feature analysis unit in this embodiment analyzes the overall image features based on the on-site grayscale image, sets the overall image features to A(i), and sets A(i)=σ3(i), where σ3(i) represents the image features of the on-site grayscale image and i represents the on-site grayscale image number. The on-site grayscale image number is defined as a number that distinguishes different on-site photo information.

[0070] Specifically, the image features of the on-site grayscale image described in this embodiment are analyzed in the same manner as the image features of the video frame image, and will not be further elaborated in this embodiment.

[0071] Please continue reading Figure 1 As shown, the multimodal target data intelligent processing system also includes:

[0072] The personnel information analysis module is used to store the associated personnel information and match the associated personnel information with preset extracted keywords to count the number of text matches and the number of matching personnel, and is also used to analyze personnel association parameters based on the number of text matches and the number of matching personnel. The personnel information analysis module is connected to the target data acquisition module.

[0073] Specifically, the personnel information analysis module in this embodiment matches the target text with preset extraction keywords, and counts the number of times the text extraction keywords appear in the target text as the number of text matches. If N1(j)>0, the personnel information analysis module sets the currently analyzed associated personnel information as a matched person; if N1(j)=0, the personnel information analysis module sets the currently analyzed associated personnel information as an unmatched person; the personnel information analysis module counts the number of associated personnel information of the matched person as the number of matched persons, and analyzes the personnel association parameter based on the number of text matches and the number of matched persons, setting the personnel association parameter to R(j), setting R(j)=N1(j)×lg[NR1 / (NR2+1)] / N2(j), where N1(j) represents the number of text matches, N2(j) represents the number of words in the target text, NR1 represents the number of associated personnel information, NR2 represents the number of matched persons, and j represents the personnel number. The personnel number is defined as a number that distinguishes different associated personnel information.

[0074] Specifically, in this embodiment, the preset extraction keywords include but are not limited to the target person's name and on-site location, etc., which only need to meet the extraction of important content in the target text. For example, keywords such as time and license plate number can also be added.

[0075] Specifically, in this embodiment, the personnel information analysis module is used to analyze the matching of the associated personnel information and the preset extracted keywords to analyze the number of text matches and the number of matching personnel, thereby analyzing the personnel association parameters, and using the personnel association parameters to represent the association characteristics between the key contents in the target text, to achieve the analysis of the word frequency characteristics in the target text, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0076] Please continue reading Figure 1 As shown, the multimodal target data intelligent processing system also includes:

[0077] The data storage management module is used to analyze the feature correlation based on the local features of the image, the overall features of the image, the on-site video image and the personnel association parameters, and store the target scene information and the associated personnel information according to the feature correlation. The data storage association module is connected to the image information analysis module and the personnel information analysis module.

[0078] See also Figure 4 As shown, the data storage management module includes:

[0079] The feature correlation analysis unit is used to analyze the feature correlation degree based on the overall image characteristics, on-site video images and personnel correlation parameters.

[0080] Specifically, the feature correlation analysis unit in this embodiment analyzes the feature correlation degree according to the overall image features, the on-site video image and the personnel correlation parameters, sets the feature correlation degree as F(m), and sets , where H1(m) represents the vector of the overall image features, H1(m)=[A(i)], H2(m) represents the vector of the image features of the on-site video image, H2(m)=[σ1(k)], H3(m) represents the vector of the personnel-related parameters, H3(m)=[R(j)], and m represents the acquisition data number. The acquisition data number is defined as the number that distinguishes the target scene information and associated personnel information of different groups.

[0081] Specifically, in this embodiment, the feature association analysis unit analyzes the overall image features, on-site video images and personnel association parameters to analyze the feature association degree, and uses the feature association degree to represent the association features between the on-site monitoring video, on-site grayscale image and associated personnel information, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0082] Please continue reading Figure 4 As shown, the data storage management module also includes:

[0083] The local feature matching unit is used to analyze the local feature correlation according to the local features of the image and the overall features of the image. The local feature matching unit is connected to the feature correlation analysis unit.

[0084] Specifically, the local feature matching unit in this embodiment uses the split number corresponding to the image split area with obvious local features as the obvious feature number set, and analyzes the local feature correlation based on the obvious feature number set, the local features of the image and the overall features of the image. If B≤b, the local feature matching unit determines that the local feature correlation is weak; otherwise, the local feature matching unit determines that the local feature correlation is strong, and processes the feature correlation analysis process. The feature correlation after processing is F1(m), and F1(m)=F(m)×e 1-B ; Where B represents the local feature matching parameter, set , U1 represents the set of distinct feature numbers, NU1 represents the number of split numbers in the set of distinct feature numbers, and b represents the local correlation threshold, where 0.7≤b≤0.9. It is understood that the value of the local correlation threshold is not specifically limited in this embodiment, and can be freely set by those skilled in the art, as long as it satisfies the analysis of local feature correlation. The optimal value of the local correlation threshold is: b = 0.8.

[0085] Specifically, in this embodiment, the local feature matching unit analyzes the local features and the overall features of the image to analyze the local feature matching parameters, and uses the local feature matching parameters to represent the difference relationship between the overall features of the on-site grayscale image and the local features of the image in the image split area where the local features are obvious, thereby analyzing the local feature correlation. When the local feature correlation is strong, the feature correlation analysis process is processed so that the processed feature correlation is related to the local features of the image, thereby increasing the accuracy of image data processing, thereby improving the system's analysis efficiency of the target data and improving the accuracy of target data storage and processing.

[0086] Please continue reading Figure 4 As shown, the data storage management module also includes:

[0087] The storage matching analysis unit is used to match the stored target scene information and related personnel information according to the feature correlation degree, and analyze the storage correlation according to the matching result. The storage matching analysis unit is connected to the local feature matching unit.

[0088] Specifically, the storage matching analysis unit in this embodiment counts the number of stored target scene information and associated person information that satisfy |F(m) / F(z)-1|≤α1 as the stored matching number, and extracts stored target scene information that satisfies |A(i) / A(z)-1|≤α1 and has the same photo type as the matching scene information. The storage matching analysis unit counts the number of matching scene information with storage type one as the first matching number, and counts the number of matching scene information with storage type two as the second matching number. Here, F(z) represents the feature correlation degree corresponding to the stored target scene information and the associated person information, z represents the storage information number, A(z) represents the overall image feature corresponding to the stored target scene information, and α1 represents the first storage matching threshold, where 0.05≤α1≤0.1. It will be understood that the value of the first storage matching threshold is not specifically limited in this embodiment and can be freely set by those skilled in the art. The optimal value of the first storage matching threshold is: α1 = 0.08.

[0089] Specifically, the storage matching analysis unit in this embodiment analyzes the storage association according to the storage matching number, the first matching number, and the second matching number. If [NA2 / (NA1+NA2)] NF / Nm >α2, the storage matching analysis unit determines that the storage association is weak; conversely, the storage matching analysis unit determines that the storage association is strong and processes the analysis process of the local feature matching parameters. The processed local feature matching parameter is B1, and B1=B×(NA1+NA2) / NA2 is set; wherein NA1 represents the first match number, NA2 represents the second match number, NF represents the storage match number, Nm represents the number of stored target site information and associated personnel information, and α2 represents the second storage matching threshold, 1≤α2≤1.2. It can be understood that the value of the second storage matching threshold is not specifically limited in this embodiment, and those skilled in the art can freely set it as long as it satisfies the analysis of storage association. The optimal value of the second storage matching threshold is α2=1.1.

[0090] Specifically, in this embodiment, the stored target scene information and related personnel information are analyzed by the storage matching analysis unit to analyze the storage matching number, the first matching number and the second matching number, and to extract and analyze similar data of the currently analyzed collected data, thereby analyzing the storage correlation. When the storage correlation is strong, the analysis process of the local feature matching parameters is processed so that the local feature matching parameters are related to the correlation between the current collected data and the stored data, thereby improving the system's analysis efficiency of the target data and improving the accuracy of the target data storage and processing.

[0091] Please continue reading Figure 4 As shown, the data storage management module also includes:

[0092] The data storage management unit is used to analyze the storage types of the target scene information and the associated personnel information according to the feature correlation degree. The data storage management unit is connected to the storage matching analysis unit.

[0093] Specifically, the data storage management unit in this embodiment analyzes the storage type of the target scene information and associated personnel information based on the feature correlation. If F(m) ≤ f, the data storage management unit determines that the storage type of the target scene information and associated personnel information is Class I; otherwise, the data storage management unit determines that the storage type of the target scene information and associated personnel information is Class II. Where f represents the storage management threshold, 0.6 ≤ f ≤ 0.9. It will be appreciated that this embodiment does not impose specific limitations on the value of the storage management threshold; those skilled in the art may freely set it, as long as it satisfies the storage type determination. The optimal value of the storage management threshold is: f = 0.8.

[0094] Specifically, in this embodiment, the storage type is one, which means that the collected target site information and related personnel information are cold data, and the currently collected data can be stored in an ordinary storage device. The storage type is two, which means that the collected target site information and related personnel information are hot data, which are key data for data analysis. The currently collected data can be stored in a high-performance storage device to improve the efficiency of subsequent data retrieval and analysis.

[0095] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not limitations on the implementation methods of the present invention. For ordinary technicians in this field, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation methods here. All obvious changes or modifications derived from the technical solution of the present invention are still within the scope of protection of the present invention.

Claims

1. A multimodal target data intelligent processing system, characterized in that: include: Target data collection module, used to collect target site information and related personnel information; A video data extraction module is used to extract the target area of ​​the video according to the target scene information, and is also used to analyze the change characteristics of the video frame according to the target scene information, and extract the scene video image according to the video frame change characteristics and the video target area; An image information analysis module is used to analyze the local features and overall features of the image based on the target scene information, and is also used to analyze the degree of local feature significance based on the local features of the image; The personnel information analysis module is used to store the associated personnel information and match the associated personnel information with the preset extracted keywords to count the number of text matches and the number of matching personnel, and also to analyze the personnel association parameters based on the number of text matches and the number of matching personnel; A data storage management module is used to analyze feature correlation based on local image features, overall image features, on-site video images, and personnel correlation parameters, and store target scene information and associated personnel information based on the feature correlation; The personnel information analysis module matches the target text with the preset extraction keywords, and counts the number of times the text extraction keywords appear in the target text as the text matching number N1(j). If N1(j)>0, the personnel information analysis module sets the current analysis associated personnel information as the matching personnel; if N1(j)=0, the personnel information analysis module sets the current analysis associated personnel information as the unmatched personnel; the personnel information analysis module counts the number of associated personnel information of the matching personnel as the matching personnel number NR2, and analyzes the personnel association parameter according to the text matching number N1(j) and the matching personnel number NR2 to obtain the personnel association parameter R(j), R(j)=N1(j)×lg[NR1 / (NR2+1)] / N2(j), where N1(j) represents the text matching number, N2(j) represents the number of words in the target text, NR1 represents the number of associated personnel information, NR2 represents the number of matching personnel, and j represents the personnel number. The data storage management module is provided with a feature correlation analysis unit, which is used to analyze the feature correlation degree according to the overall image feature A(i), the on-site video image and the personnel correlation parameter R(j) to obtain the feature correlation degree F(m). ,where H1(m) represents the vector of the overall image features, H1(m)=[A(i)], H2(m) represents the vector of the image features of the on-site video image, H2(m)=[σ1(k)], H3(m) represents the vector of personnel-related parameters, H3(m)=[R(j)], and m represents the number of the collected data; The data storage management module is further provided with a local feature matching unit, which is used to use the split number corresponding to the image split area with obvious local features as the obvious feature number set U1, and analyze the local feature correlation according to the obvious feature number set U1, the local feature W(v) of the image and the overall feature A(i) of the image. The local feature correlation includes weak local feature correlation and strong local feature correlation, and processes the feature correlation analysis process when the local feature correlation is strong. The feature correlation after processing is F1(m), F1(m)=F(m)×e1-B; wherein B represents the local feature matching parameter, and the setting , U1 represents the set of obvious feature numbers, NU1 represents the number of split numbers in the set of obvious feature numbers; The data storage management module is further provided with a storage matching analysis unit, which is used to count the number of stored target scene information and associated personnel information that satisfy |F(m) / F(z)-1|≤α1 as the storage matching number NF, and extract the stored target scene information that satisfies |A(i) / A(z)-1|≤α1 and has the same photo type as the matching scene information, and the storage matching analysis unit counts the number of matching scene information of the storage type of the first type as the first matching number NA1, and counts the number of matching scene information of the storage type of the second type as the second matching number NA2, wherein F(z) represents the feature correlation degree corresponding to the stored target scene information and the associated personnel information, z represents the storage information number, A(z) represents the overall image feature corresponding to the stored target scene information, and α1 represents the first storage matching threshold; The storage matching analysis unit analyzes the storage correlation according to the storage matching number NF, the first matching number NA1 and the second matching number NA2. The storage correlation includes strong storage correlation and weak storage correlation. When the storage correlation is strong, the analysis process of the local feature matching parameters is processed. The processed local feature matching parameters are B1, B1=B×(NA1+NA2) / NA2.

2. The multimodal target data intelligent processing system according to claim 1, characterized in that: The video data extraction module is provided with a change feature analysis unit, which is used to extract an image frame every K frames from the on-site monitoring video as a video frame image, and number the video frame images according to the extraction order to obtain a video frame number, where K represents a video extraction parameter; The change feature analysis unit analyzes the video frame change feature according to the video frame image to obtain the video frame change feature Q(k), Q(k)=σ1(k) / [σ1(k-1)+1], wherein σ1(k) represents the image feature of the video frame image, the image feature of the video frame image is the standard deviation of the grayscale values ​​of the pixels in the video frame image, and k represents the video frame number.

3. The multimodal target data intelligent processing system according to claim 2, characterized in that: The video data extraction module is also provided with a video image extraction unit, which is used to extract the on-site video image according to the video frame change feature Q(k). If there is a target area in the on-site monitoring video and Q(k)≤q, the video image extraction unit currently analyzes the video frame image corresponding to the video frame change feature as the on-site video image; otherwise, the video image extraction unit does not extract the on-site video image; wherein q represents the change feature threshold.

4. The multimodal target data intelligent processing system according to claim 3, characterized in that: The image information analysis module is provided with a division region analysis unit, which is used to analyze the local features of the image according to the image split region to obtain the local image features W(v), W(v)=σ2(v), where σ2(v) represents the image features of the image split region and v represents the split number; The divided region analysis unit analyzes the local feature obviousness according to the local feature W(v) of the image, and the local feature obviousness includes the local feature not obvious and the local feature obvious.

5. The multimodal target data intelligent processing system according to claim 4, characterized in that: The image information analysis module is also provided with a picture feature analysis unit, which is used to analyze the overall image features based on the on-site grayscale image to obtain the overall image features A(i), A(i)=σ3(i), where σ3(i) represents the image features of the on-site grayscale image and i represents the on-site grayscale image number.

6. The multimodal target data intelligent processing system according to claim 1, characterized in that: The data storage management module is further provided with a data storage management unit, which is used to analyze the storage types of the target scene information and the associated personnel information according to the feature correlation degree, and the storage types include the first category and the second category.

Citation Information

Patent Citations

  • Multi-modal retrieval method and device and storage medium

    CN117763174A

  • Image analysis function based video monitoring system and storage method thereof

    CN104219491A

  • Attribute recognition-based popular science content matching system

    CN118606513A