Hierarchical fusion method and device based on cross-media understanding technology and related media
By employing a layered fusion approach based on cross-media understanding technology, and comprehensively utilizing image, video, audio, text, and sensor information, the limitations of single sensing devices in public safety monitoring are addressed, enabling comprehensive information collection and efficient anomaly detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2026-03-27
AI Technical Summary
In public safety monitoring, the use of a single sensing device leads to limitations in the detection and identification of abnormal events and low efficiency in information acquisition.
A layered fusion method using cross-media understanding technology is adopted to acquire image, video, audio, text and sensor information, perform data feature recognition and feedback information fusion, establish a multi-source information fusion system, identify abnormal events and trigger alarms.
It improves the efficiency of information acquisition and the accuracy of abnormal event detection and identification, makes up for the limitations of a single sensing device, and realizes comprehensive information collection and abnormal event detection.
Smart Images

Figure CN115828184B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent information fusion, and in particular to a layered fusion method and device based on cross-media understanding technology and related media. BACKGROUND
[0002] At present, single image, video, sound, text and other information understanding technology has been relatively mature, and is applied in many places; but using a single sensing means in the process of public safety monitoring cannot achieve all-around sensing and all-around information collection; for example, the image information cannot be obtained when the pedestrian target in the image and video appears to be blocked or for special occasions such as toilets; the pedestrian target detection effect is not ideal when the night scene image quality is poor; for example, the sound recognition effect is greatly reduced when the sound is in a noisy environment or the sound source is far away from the sound collector; in addition, it is difficult to obtain information by text detection and recognition, and special equipment needs to be used to assist in obtaining. Therefore, using a single sensing device cannot achieve all-around sensing and all-around information collection, and there is limitation in the detection and identification of abnormal events and low information acquisition efficiency. SUMMARY
[0003] The embodiment of the present application provides a layered fusion method, device and related media based on cross-media understanding technology, aiming to solve the problem of low information acquisition efficiency and limitation in the detection and identification of abnormal events caused by using a single sensing device in the process of public safety monitoring.
[0004] In a first aspect, the embodiment of the present application provides a layered fusion method based on cross-media understanding technology, comprising:
[0005] obtaining data information, taking the data information as a data layer; wherein the data information includes image information, video information, audio information, text information and sensor information;
[0006] detecting data features according to the data information, taking the data features as a feature layer; wherein the data features include image features, video features, audio features, text features and sensor features;
[0007] recognizing feature recognition results according to the data features, taking the feature recognition results as a decision layer; wherein the feature recognition results include image feature recognition results, video feature recognition results, audio feature recognition results, text feature recognition results and sensor feature recognition results;
[0008] obtaining feedback information, taking the feedback information as a feedback layer; wherein the feedback information includes image feedback information, video feedback information, audio feedback information, text feedback information and sensor feedback information;
[0009] fuse the data layer, the feature layer, the decision layer and the feedback layer to obtain a multi-source information fusion result;
[0010] determine whether an abnormal event occurs according to the multi-source information fusion result; if not, information is re-acquired and determination is performed again; if yes, an alarm is triggered and on-site personnel feedback information is acquired, and the on-site personnel feedback information is transmitted into the feedback layer.
[0011] In a second aspect, an embodiment of the present application provides a layered fusion device based on cross-media understanding technology, comprising:
[0012] an information acquisition unit, configured to acquire data information and take the data information as a data layer; wherein the data information comprises image information, video information, audio information, text information and sensor information;
[0013] a feature calculation unit, configured to detect data features according to the data information and take the data features as a feature layer; wherein the data features comprise image features, video features, audio features, text features and sensor features;
[0014] a feature recognition unit, configured to recognize feature recognition results according to the data features and take the feature recognition results as a decision layer; wherein the feature recognition results comprise image feature recognition results, video feature recognition results, audio feature recognition results, text feature recognition results and sensor feature recognition results;
[0015] a feedback acquisition unit, configured to acquire feedback information and take the feedback information as a feedback layer; wherein the feedback information comprises image feedback information, video feedback information, audio feedback information, text feedback information and sensor feedback information;
[0016] an information fusion unit, configured to fuse the data layer, the feature layer, the decision layer and the feedback layer to obtain a multi-source information fusion result;
[0017] a result determination unit, configured to determine whether an abnormal event occurs according to the multi-source information fusion result; if not, information is re-acquired and determination is performed again; if yes, an alarm is triggered and on-site personnel feedback information is acquired, and the on-site personnel feedback information is transmitted into the feedback layer.
[0018] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor and a computer program stored in the memory and capable of running on the processor, and the processor implements the layered fusion method based on cross-media understanding technology of the first aspect when executing the computer program.
[0019] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the hierarchical fusion method based on cross-media understanding technology of the first aspect.
[0020] The embodiment of the present application provides a hierarchical fusion method based on cross-media understanding technology, acquires data information, takes the data information as a data layer, detects data features according to the data information, takes the data features as a feature layer, identifies feature recognition results according to the data features, takes the feature recognition results as a decision layer, acquires feedback information, takes the feedback information as a feedback layer, fuses the data layer, the feature layer, the decision layer and the feedback layer to obtain a multi-source information fusion result, judges whether an abnormal event occurs according to the multi-source information fusion result, if not, reacquires information and judges again, if yes, triggers an alarm and acquires on-site personnel feedback information, and transmits the on-site personnel feedback information into the feedback layer. The present application identifies abnormal events by fusing multi-source information, establishes a cross-media detection and identification system, and thus improves information acquisition efficiency and abnormal event detection and identification precision.
[0021] The embodiment of the present application also provides a hierarchical fusion device based on cross-media understanding technology, a computer device and a storage medium, which also have the beneficial effects described above. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0023] Figure 1 A flowchart of a hierarchical fusion method based on cross-media understanding technology provided by the embodiment of the present application;
[0024] Figure 2 A network architecture diagram of a hierarchical fusion method based on cross-media understanding technology provided by the embodiment of the present application;
[0025] Figure 3 A schematic block diagram of a hierarchical fusion device based on cross-media understanding technology provided by the embodiment of the present application. DETAILED DESCRIPTION
[0026] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are some of the embodiments of the present application but not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the scope of the present application.
[0027] It should be understood that the terms "comprising" and "including" as used in the specification and the appended claims indicate the presence of the described features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0028] It should also be understood that the terms used in the present application specification are only for the purpose of describing particular embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0029] It should be further understood that the term "and / or" as used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations thereof, and includes these combinations.
[0030] Please see the following Figure 1 , Figure 1 A flowchart of a hierarchical fusion method based on a cross-media understanding technology according to an embodiment of the present application is shown in the figure, which specifically includes steps S101-S106.
[0031] S101, data information is acquired, and the data information is taken as a data layer; the data information includes image information, video information, audio information, text information and sensor information;
[0032] S102, data features are detected according to the data information, and the data features are taken as a feature layer; the data features include image features, video features, audio features, text features and sensor features;
[0033] S103, feature recognition results are recognized according to the data features, and the feature recognition results are taken as a decision layer; the feature recognition results include image feature recognition results, video feature recognition results, audio feature recognition results, text feature recognition results and sensor feature recognition results;
[0034] S104, acquire feedback information, and take the feedback information as a feedback layer; wherein the feedback information includes image feedback information, video feedback information, audio feedback information, text feedback information, and sensor feedback information;
[0035] S105, fuse the data layer, the feature layer, the decision layer, and the feedback layer to obtain a multi-source information fusion result;
[0036] S106, determine whether an abnormal event occurs according to the multi-source information fusion result; if not, reacquire information and determine again; if yes, trigger an alarm and acquire on-site personnel feedback information, and input the on-site personnel feedback information into the feedback layer.
[0037] In combination Figure 1 and Figure 2 As shown in FIGS. 1 to 3, in step S101, first, data information needs to be acquired, and the data information includes image information, video information, audio information, text information, and sensor information; the image information and the video information can be acquired by a monitoring camera, and the camera generally adopts a combination of a gun camera and a ball camera to capture pictures in linkage, the gun camera has a better image capture effect, and the ball camera can automatically zoom and can control to capture a specific picture; the audio information can be acquired by a sound pickup or other hardware capable of capturing a sound source; the text information is text information on an electronic screen (a screen of a mobile phone, a tablet, a computer, etc.) captured by the camera, and in an application process, all information on the electronic screen can be extracted to obtain the text information for subsequent analysis; it should be noted that the sensor information here is not only one type of information, and the sensor information can be smoke sensor information, flame sensor information, odor sensor information, laser sensor information, etc., different sensors can be added here according to actual conditions to capture different sensor information; of course, the more types of sensors are set, the more comprehensive the information capture is, and the more accurate the subsequent information analysis is; finally, after the image information, the video information, the audio information, the text information, and the sensor information are acquired, all the information is summarized as a data layer, that is, the data information is taken as a data layer.
[0038] In an embodiment, the step S101 includes:
[0039] Whether the acquired data information is useful information is determined according to the following formula:
[0040] SNR = 10lg (Ps / Pn)
[0041] wherein Ps represents signal effective power; Pn represents noise effective power; and SNR represents signal-to-noise ratio.
[0042]
[0043] wherein, p signal represents the useful signal power; p noise +p distortion represents the useless signal power; SINAD represents the signal-to-noise ratio;
[0044] screening the useful information and eliminating the useless information.
[0045] In the embodiment, when one signal source cannot provide useful information, other signal sources can provide auxiliary information, ensuring information acquisition in time and adaptive adjustment; specifically, whether a signal source can provide useful information can be determined by SNR (signal-to-noise ratio, the same below) and SINAD (signal-to-noise ratio, the same below); for example, when the sound signal-to-noise ratio is lower than a certain threshold, it means that the noise mixed in the audio signal is too large, and the audio signal cannot provide useful information; similarly, the video information, the image information and the sensor information also have SNR and SINAD, useful information can be screened and useless information can be eliminated, so as to improve the efficiency of obtaining the data information.
[0046] In an embodiment, the step S101 further comprises:
[0047] The weight coefficients of the image information, the video information, the audio information, the text information and the sensor information are respectively calculated according to the following formula:
[0048]
[0049] wherein, represents the weight coefficient of the i-th information; SNR i represents the signal-to-noise ratio of the i-th information; SNR img represents the signal-to-noise ratio of the image information; SNR audio represents the signal-to-noise ratio of the audio information; SNR text represents the signal-to-noise ratio of the text information; SNR other represents the signal-to-noise ratio of the sensor information; SNR vldeo represents the signal-to-noise ratio of the video information.
[0050] In the embodiment, the acquisition of the image information, the video information, the audio information, the text information and the sensor information adds an attention mechanism, such as increasing the weight coefficient of acquiring the image information in the daytime (the image is clearer in the daytime with sufficient light), increasing the weight coefficient of acquiring the audio information at night (there are fewer pedestrians and vehicles and less environmental noise at night); the weight coefficients of other information can also be adjusted according to actual conditions, and the weight coefficients here can be flexibly changed; it should be noted that i in the formula can represent image, audio, text or other sensor information.
[0051] In an embodiment, the step S101 further comprises:
[0052] The data information is calculated according to the following formula:
[0053]
[0054] wherein, Da img represents the image information; Da video represents the video information; Da audto represents the audio information; Da text represents the text information; Da other represents the sensor information; S 数据 represents the data information; represents the image weight; represents the video weight; represents the audio weight; represents the text weight; represents the sensor weight.
[0055] In the embodiment, the data information is obtained by multiplying the image information, the video information, the audio information, the text information and the sensor information by corresponding weight coefficients and then adding them; wherein the sensor information includes various sensor information, such as smoke sensor information, flame sensor information, etc.; it should be noted that S 数据 represents the data information, which can also be regarded as the data layer; the image information can be the acquisition of current scene pedestrian information and related motion information, for example, Da img = 0 can represent no one, i.e. no information of any pedestrian is acquired, Da img = 1, which indicates that the information of pedestrians is acquired; of course, Da img the acquisition of vehicles and other target information is also applicable, represents the image weight, in combination with Da imgThe image information with specified weight can be acquired; the audio information can be the target audio after removing background sound after acquiring sound, or all audio directly acquired according to actual conditions; similarly, the text information, the video information and the sensor information are acquired in combination with corresponding weights, and the acquisition manner of the image information and the audio information can be referred to.
[0056] In step S102, data features are detected according to the data information; the data features include image features, video features, audio features, text features and sensor features; after the data information is acquired, the data information is subjected to feature detection to obtain corresponding features; after the image features, the video features, the audio features, the text features and the sensor features are acquired, all acquired features are fused, that is, the data features are taken as a feature layer.
[0057] In an embodiment, the step S102 includes:
[0058] The data features are calculated according to the following formula:
[0059]
[0060] Fei represents the image features; Fev represents the video features; Fea represents the audio features; Fext represents the text features; Fes represents the sensor features; and S represents the data features. img video audio text other 特征
[0061] In the embodiment, the data features are obtained by adding the image features, the video features, the audio features, the text features and the sensor features after being multiplied by corresponding weight coefficients respectively; the detection of the image features and the video features, such as whether the pedestrian has pursuit, fight, whether to hold a stick or a knife, and other features (the pursuit of the pedestrian in the time sequence, the action information, the fight can refer to the openpose-human pose estimation algorithm, and the limb key point features, the stick or knife features); if the camera detects abnormal behavior, the gun and the ball machine are linked to enlarge and capture, at this time, if the face is detected, the face features are extracted, and then stored and compared; here, the storage and comparison are the same as the comparison with the pre-recorded feature information, and the same is true for the following; the detection of the audio features, such as audio content, keywords, voiceprint features, and the like (speech recognition, audio conversion text recognition, voiceprint recognition are realized by feature extraction and comparison, and there are mature algorithms GMM, DNN, based on attention mechanism, Learning to rank and other algorithms); if the keyword information is detected in the audio signal, the camera and the sound pickup are linked to enlarge to the sound source position to capture, the face and voiceprint features are extracted, and stored and compared; the detection of the text features can detect text content, keywords and the like; if the keyword information is detected in the text signal, the camera is linked to enlarge to the text signal source position to capture, the face is extracted, and stored and compared; for the detection of the sensor features, such as the characteristic that the laser sensor is the infrared radiation generated by the object; in addition, the weight coefficient calculation of the data features can refer to the weight coefficient setting method of the data information.
[0062] In step S103, a feature recognition result is obtained according to the data features, and the feature recognition result is used as a decision layer; the feature recognition result includes an image feature recognition result, a video feature recognition result, an audio feature recognition result, a text feature recognition result and a sensor feature recognition result; it should be noted that the decision layer is used to decide whether to perform an alarm operation according to the feature recognition result.
[0063] In an embodiment, the step S103 includes:
[0064] The feature recognition result is calculated according to the following formula:
[0065]
[0066] wherein, De img represents the image feature recognition result; De video represents the video feature recognition result; De audio represents the audio feature recognition result; De text represents the text feature recognition result; De otherS represents the sensor feature recognition result; S 决策 S represents the feature recognition result.
[0067] In the embodiment, the feature recognition result is obtained by adding the image feature recognition result, the video feature recognition result, the audio feature recognition result, the text feature recognition result and the sensor feature recognition result after being multiplied by corresponding weight coefficients; it is to be noted that S 决策 S represents the feature recognition result, which is equivalent to the decision layer; in addition, whether to alarm is decided according to the image feature recognition result (such as alarm: 1, no alarm: 0, which can be adjusted); similarly, whether to alarm is decided according to the audio feature recognition result, the video feature recognition result, the text feature recognition result and the sensor feature recognition result respectively; the weight coefficient calculation of the feature recognition result can also refer to the weight coefficient setting method of the data information.
[0068] In step S104, feedback information is acquired, which is feedback information received by monitoring personnel or on-site investigation personnel; the feedback information includes image feedback information, video feedback information, audio feedback information, text feedback information and sensor feedback information; after the feedback information is acquired, the feedback information is taken as a feedback layer.
[0069] In an embodiment, the step S104 includes:
[0070] The feedback information is calculated according to the following formula:
[0071]
[0072] Wherein, Fb img represents the image feedback information; Fb video represents the video feedback information; Fb audio represents the audio feedback information; Fb text represents the text feedback information; Fb other represents the sensor feedback information; S 反馈 S represents the feedback information.
[0073] In the embodiment, the feedback information is obtained by adding the image feedback information, the video feedback information, the audio feedback information, the text feedback information and the sensor feedback information after being multiplied by corresponding weight coefficients; the image feedback information can be the feedback of whether to alarm after the monitoring personnel sees the image, and then the feedback information of the monitoring personnel is obtained (for example, alarm: 1, no alarm: 0, which can be adjusted); similarly, for the audio feedback information, the video feedback information, the text feedback information and the sensor feedback information, the monitoring personnel can also give feedback according to the corresponding information; the weight coefficient calculation of the feedback information can also refer to the weight coefficient setting method of the data information.
[0074] In step S105, the data layer, the feature layer, the decision layer and the feedback layer are fused to obtain a multi-source information fusion result; by establishing a layered multi-source data fusion mechanism, the limitations of any single layer information fusion are made up; the data layer has large data fusion amount but less information loss; the feature layer is affected by feature extraction, and compared with the data layer fusion, the data amount is reduced but the information loss is large; the decision layer is the recognition result of multi-source data, and has no requirement for the consistency of multi-source data, has the highest fusion efficiency, and is the most common way of multi-source information fusion, but also has the largest information loss, and when the data quality of a single source information is poor, the recognition result is wrong, which leads to that this kind of data has a great influence on the final result; the feedback layer belongs to the alarm information confirmation mechanism of the public security monitoring personnel, and can be used for model self-adaptive learning adjustment.
[0075] In step S106, it is judged whether an abnormal event occurs according to the multi-source information fusion result; here, the judgment of the abnormal event is actually a linkage feedback of the judgment results of the image, audio, text and sensor; in this way, compared with a single sensor signal source, the multi-source information perception and information collection are more comprehensive; for the single-layer fusion mechanism in the multi-source signal fusion, the layered fusion result is more accurate; at the same time, the linkage feedback mechanism is increased, the system self-adapts learning through the feedback supervision of the ball machine and the gun machine, the camera and the pickup, the monitoring personnel and the alarm signal, so that the next judgment of the abnormal event is more accurate; when it is judged that no abnormal event occurs, the information of the image, video, audio, text and sensor is reacquired and judged again; when it is judged that an abnormal event occurs, an alarm is triggered and the on-site personnel feedback information is acquired, and the on-site personnel feedback information is transmitted into the feedback layer; it should be noted that the feedback information is an interactive process of a system and a user, and the information seen and heard by the monitoring personnel may have lag and incomplete information, so the feedback may have deviation; of course, the monitoring personnel may not feedback, and the feedback content is empty at this time; through the information feedback by the monitoring personnel, the self-learning ability of the system can be continuously improved, and the feedback results of the monitoring personnel are correct in most cases, so the system can take the information feedback results of the monitoring personnel as training data to retrain the model and continuously improve the detection and recognition precision; through the comprehensive judgment of the abnormal event by fusing the multi-source information, the information is fused and reused in a higher dimensional level, the comprehensive perception ability is established, the limitations of a single signal are made up, and the information acquisition efficiency and the abnormal event detection and recognition precision are greatly improved.
[0076] In combination Figure 3 as shown, Figure 3 A schematic block diagram of a layered fusion device based on cross-media understanding technology provided by the embodiment of the present application is shown, and the layered fusion device based on cross-media understanding technology 300 comprises:
[0077] An information acquisition unit 301 is configured to acquire data information, and the data information is taken as a data layer; wherein the data information comprises image information, video information, audio information, text information and sensor information;
[0078] A feature calculation unit 302 is configured to detect data features according to the data information, and the data features are taken as a feature layer; wherein the data features comprise image features, video features, audio features, text features and sensor features;
[0079] A feature recognition unit 303 is configured to recognize feature recognition results according to the data features, and the feature recognition results are taken as a decision layer; wherein the feature recognition results comprise image feature recognition results, video feature recognition results, audio feature recognition results, text feature recognition results and sensor feature recognition results;
[0080] The feedback acquisition unit 304 is configured to acquire feedback information, and the feedback information is taken as a feedback layer; wherein the feedback information includes image feedback information, video feedback information, audio feedback information, text feedback information and sensor feedback information.
[0081] The information fusion unit 305 is configured to fuse the data layer, the feature layer, the decision layer and the feedback layer to obtain a multi-source information fusion result.
[0082] The result judgment unit 306 is configured to judge whether an abnormal event occurs according to the multi-source information fusion result; if not, information is re-acquired and the judgment is performed again; if yes, an alarm is triggered and on-site personnel feedback information is acquired, and the on-site personnel feedback information is transmitted to the feedback layer.
[0083] In the embodiment, first, the information acquisition unit 301 acquires data information, and the data information is taken as a data layer; the feature calculation unit 302 detects data features according to the data information, and the data features are taken as a feature layer; the feature recognition unit 303 recognizes feature recognition results according to the data features, and the feature recognition results are taken as a decision layer; the feedback acquisition unit 304 acquires feedback information, and the feedback information is taken as a feedback layer; the information fusion unit 305 fuses the data layer, the feature layer, the decision layer and the feedback layer to obtain a multi-source information fusion result; finally, the result judgment unit 306 judges whether an abnormal event occurs according to the multi-source information fusion result; if not, information is re-acquired and the judgment is performed again; if yes, an alarm is triggered and on-site personnel feedback information is acquired, and the on-site personnel feedback information is transmitted to the feedback layer.
[0084] In an embodiment, the information acquisition unit 301 is configured to:
[0085] whether the acquired data information is useful information is judged according to the following formula:
[0086] SNR=10lg(Ps / Pn)
[0087] wherein Ps represents signal effective power; Pn represents noise effective power; and SNR represents signal-to-noise ratio.
[0088]
[0089] wherein p signal represents useful signal power; p noise +p distortion represents useless signal power; and SINAD represents signal-to-noise ratio.
[0090] The useful information is screened and the useless information is eliminated.
[0091] Further, the weight coefficients of the image information, the video information, the audio information, the text information and the sensor information are calculated respectively according to the following formulas:
[0092]
[0093] wherein, represents the weight coefficient of the i-th information; SNR i represents the signal-to-noise ratio of the i-th information; SNR img represents the signal-to-noise ratio of the image information; SNR audio represents the signal-to-noise ratio of the audio information; SNR text represents the signal-to-noise ratio of the text information; SNR other represents the signal-to-noise ratio of the sensor information; SNR video represents the signal-to-noise ratio of the video information.
[0094] Further, the data information is calculated according to the following formula:
[0095]
[0096] wherein, Da img represents the image information; Da video represents the video information; Da audio represents the audio information; Da text represents the text information; Da other represents the sensor information; S 数据 represents the data information; represents the image weight; represents the video weight; represents the audio weight; represents the text weight; represents the sensor weight.
[0097] In an embodiment, the feature calculation unit 302 is configured to:
[0098] The data feature is calculated according to the following formula:
[0099]
[0100] wherein, De img represents the image feature; Fe video represents the video feature; Fe audio represents the audio feature; Fe text represents the text feature; Fe other represents the sensor feature; S 特征representing the data feature.
[0101] In an embodiment, the feature identification unit 303 is configured to:
[0102] The feature identification result is calculated according to the following formula:
[0103]
[0104] wherein De img represents the image feature identification result; De video represents the video feature identification result; De audio represents the audio feature identification result; De text represents the text feature identification result; De other represents the sensor feature identification result; S 决策 represents the feature identification result.
[0105] In an embodiment, the feedback acquisition unit 304 is configured to:
[0106] The feedback information is calculated according to the following formula:
[0107]
[0108] wherein Fb img represents the image feedback information; Fb video represents the video feedback information; Fb audio represents the audio feedback information; Fb text represents the text feedback information; Fb other represents the sensor feedback information; S 反馈 represents the feedback information.
[0109] Since the embodiments of the device part correspond to the embodiments of the method part, the embodiments of the device part are described in the description of the embodiments of the method part, and are not described here in detail.
[0110] The embodiments of the present application also provide a computer readable storage medium, which has a computer program stored thereon, and the computer program is executed to implement the steps provided by the above embodiments. The storage medium can include: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0111] The embodiment of the present application further provides a computer device, which can comprise a memory and a processor, the memory has a computer program stored therein, and the processor can realize the steps provided by the above embodiment when calling the computer program in the memory. Of course, the computer device can further comprise various network interfaces, power supplies and other components.
[0112] The various embodiments are described in the specification by way of progressive progression, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be mutually referred to. For the system disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part. It should be noted that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0113] It should be further noted that, in the specification, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
Claims
1. A layered fusion method based on cross-media understanding technology, characterized in that, include: Acquire data information and use the data information as a data layer; wherein the data information includes: image information, video information, audio information, text information, and sensor information; Data features are detected based on the data information, and these data features are used as a feature layer; wherein, the data features include: image features, video features, audio features, text features, and sensor features; The feature recognition results are obtained based on the data feature recognition, and the feature recognition results are used as the decision layer; wherein, the feature recognition results include: image feature recognition results, video feature recognition results, audio feature recognition results, text feature recognition results, and sensor feature recognition results; Obtain feedback information and use the feedback information as a feedback layer; wherein, the feedback information includes: image feedback information, video feedback information, audio feedback information, text feedback information, and sensor feedback information; The data layer, the feature layer, the decision layer, and the feedback layer are fused to obtain a multi-source information fusion result; Based on the results of the multi-source information fusion, determine whether an abnormal event has occurred; if not, reacquire information and determine again; if yes, trigger an alarm and acquire feedback information from on-site personnel, and transmit the on-site personnel feedback information to the feedback layer. The process of acquiring data information, using the data information as a data layer, includes: determining whether the acquired data information is useful information according to the following formula: SNR = 10 lg ( Ps / Pn ) in, Ps Indicates the effective power of the signal; Pn Indicates the effective power of the noise; SNR Indicates the signal-to-noise ratio; SINAD = in, Indicates the useful signal power; Indicates useless signal power; SINAD Indicate the information-to-satisfaction ratio; filter the useful information and remove the useless information; The weighting coefficients for the image information, video information, audio information, text information, and sensor information are calculated using the following formulas: in, Indicates the first i Weighting coefficients for various types of information; Indicates the first i The signal-to-noise ratio of this type of information; This represents the signal-to-noise ratio of the image information; This indicates the signal-to-noise ratio of the audio information; This indicates the signal-to-noise ratio of the text information; This indicates the signal-to-noise ratio of the sensor information; This indicates the signal-to-noise ratio of the video information; The data information is calculated using the following formula: S 数据 = in, This represents the image information; This refers to the video information; This refers to the audio information; This represents the text information; This indicates the sensor information; S 数据 This represents the data information; Represents image weights; Indicates video weight; Indicates audio weights; Indicates text weight; This represents the sensor weights.
2. The layered fusion method based on cross-media understanding technology according to claim 1, characterized in that, The step of detecting data features based on the data information and using the data features as a feature layer includes: The data features are calculated using the following formula: S 特征 = in, Represents the image features; This represents the video features; Indicates the audio features; This represents the text feature; Indicates the characteristics of the sensor; S 特征 This refers to the data characteristics.
3. The layered fusion method based on cross-media understanding technology according to claim 1, characterized in that, The step of obtaining feature recognition results based on the data features and using the feature recognition results as a decision layer includes: The feature recognition result is calculated using the following formula: S 决策 = in, This indicates the image feature recognition result; This indicates the video feature recognition result; This indicates the audio feature recognition result; This indicates the text feature recognition result; This indicates the sensor feature recognition result; S 决策 This indicates the feature recognition result.
4. The layered fusion method based on cross-media understanding technology according to claim 1, characterized in that, The step of obtaining feedback information, using the feedback information as a feedback layer, includes: The feedback information is calculated using the following formula: S 反馈 = in, This represents the image feedback information; This indicates the video feedback information; This indicates the audio feedback information; This indicates the text feedback information; This indicates the feedback information from the sensor; S 反馈 This indicates the feedback information.
5. A layered fusion device based on cross-media understanding technology, characterized in that, include: An information acquisition unit is used to acquire data information and use the data information as a data layer; wherein, the data information includes: image information, video information, audio information, text information, and sensor information; The feature calculation unit is used to detect data features based on the data information and use the data features as a feature layer; wherein, the data features include: image features, video features, audio features, text features, and sensor features; A feature recognition unit is used to obtain feature recognition results based on the data features, and to use the feature recognition results as a decision layer; wherein, the feature recognition results include: image feature recognition results, video feature recognition results, audio feature recognition results, text feature recognition results, and sensor feature recognition results; A feedback acquisition unit is used to acquire feedback information and use the feedback information as a feedback layer; wherein, the feedback information includes: image feedback information, video feedback information, audio feedback information, text feedback information, and sensor feedback information; An information fusion unit is used to fuse the data layer, the feature layer, the decision layer, and the feedback layer to obtain a multi-source information fusion result; The result judgment unit is used to determine whether an abnormal event has occurred based on the multi-source information fusion result; if not, it reacquires the information and judges again; if so, it triggers an alarm and acquires on-site personnel feedback information, and transmits the on-site personnel feedback information to the feedback layer. The information acquisition unit is specifically used to determine whether the acquired data information is useful information according to the following formula: SNR = 10 lg ( Ps / Pn ) in, Ps Indicates the effective power of the signal; Pn Indicates the effective power of the noise; SNR Indicates the signal-to-noise ratio; SINAD = in, Indicates the useful signal power; Indicates useless signal power; SINAD Indicate the information-to-satisfaction ratio; filter the useful information and remove the useless information; The weighting coefficients for the image information, video information, audio information, text information, and sensor information are calculated using the following formulas: in, Indicates the first i Weighting coefficients for various types of information; Indicates the first i The signal-to-noise ratio of this type of information; This represents the signal-to-noise ratio of the image information; This indicates the signal-to-noise ratio of the audio information; This indicates the signal-to-noise ratio of the text information; This indicates the signal-to-noise ratio of the sensor information; This indicates the signal-to-noise ratio of the video information; The data information is calculated using the following formula: S 数据 = in, This represents the image information; This refers to the video information; This refers to the audio information; This represents the text information; This indicates the sensor information; S 数据 This represents the data information; Represents image weights; Indicates video weight; Indicates audio weights; Indicates text weight; This represents the sensor weights.
6. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the layered fusion method based on cross-media understanding technology as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the layered fusion method based on cross-media understanding technology as described in any one of claims 1 to 4.
Citation Information
Patent Citations
A unified fusion system of government affairs data
CN109360136A
Weight-adaptive power equipment external insulation acousto-optic collaborative diagnosis method
CN115291055A