An audio-visual multi-modal recognition method and system
Through the audio and video multimodal recognition system, the coordinated work of the perception layer, identification layer and indication layer is solved, and the problem that recorders in the prior art cannot effectively assist in target tracking in law enforcement scenarios is achieved, achieving more efficient and accurate target tracking assistance effects.
Patent Information
- Application Number
- CN202510230366.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
When the prior art is used for target tracking in law enforcement scenarios, the recorder cannot effectively assist staff, resulting in target tracking being easily lost.
Design an audio and video multi-modal recognition system, through the coordinated work of the perception layer, identification layer and indication layer, lock and track targets in real time, provide continuous indications and labels, and assist staff in target tracking.
It effectively improves the functionality and purpose of the recorder in law enforcement scenarios, ensures the smooth progress of the target tracking task, and reduces the risk of target loss.
Smart Images

Figure CN119723428B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to an audio-visual multi-modal recognition method and system. Background Art
[0002] Cameras are often called recorders in daily life recordings and law enforcement scenarios. They continuously record video data in the first-person perspective for subsequent retrieval and viewing. In law enforcement scenarios, recorders are often worn on the chest or the surface of helmets for recording.
[0003] A Chinese invention patent with the application number 202410768593.X discloses a multi-modal personality trait recognition method based on audio and video. The method includes: obtaining a video to be detected and the audio data corresponding to the video to be detected; uniformly cropping the video to be detected, and extracting video frames from each cropped segment; respectively processing the extracted video frames through a face branch and a scene branch to obtain face features and scene features; preprocessing the audio data corresponding to the video to be detected, then performing a windowing operation, converting each window into frequency domain information using Fourier transform, and then using an audio processing network to process to obtain audio features; fusing the obtained face features, scene features, and audio features, inputting them into a decoder for temporal modeling, and then processing the result vector through a classifier based on a multi-layer perceptron to obtain the personality trait recognition result of the video to be detected.
[0004] This application aims to solve the problem of "the existing technology has a poor effect of integrating information between audio and visual modalities and cannot obtain more accurate personality trait recognition results".
[0005] However, existing recorders are often only used for recording video images. In the target tracking scenario of law enforcement scenarios, recorders cannot provide more auxiliary effects and completely rely on staff to chase. In a complex environment with multiple people, it is extremely easy to lose the tracking target.
[0006] Therefore, an audio-visual multi-modal recognition system is proposed. Summary of the Invention
[0007] In view of the above-mentioned drawbacks of the existing technology, the present invention provides an audio-visual multi-modal recognition method and system, which solves the technical problems proposed in the above background art.
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0009] In a first aspect, an audio-visual multi-modal recognition system includes: a perception layer, a recognition layer, and an indication layer;
[0010] The video data collected by the collection device is uploaded to the perception layer in real time. The perception layer synchronously senses the posture of the collection device, locks and tracks the target based on the collected video and the posture of the collection device, and synchronously extracts the characteristic parameters of the tracked target. The recognition layer further receives the characteristic parameters of the tracked target extracted in the perception layer, synchronously controls the perception layer to run again to lock and track the target, and continuously recognizes the tracked target in combination with the re-locked tracked target and the characteristic parameters of the tracked target. The indication layer receives in real time the recognition result of the tracked target operated by the recognition layer, and continuously indicates the tracked target based on the recognition result;
[0011] The recognition layer includes a receiving module, a control module and an identification module. The receiving module is used to receive the characteristic parameters of the tracked target extracted by the operation of the perception layer. The control module is used to control the perception layer to run again and re-lock the tracked target. The identification module is used to receive video data and continuously identify the tracked target in the video data;
[0012] The continuous recognition logic of the tracked target in the identification module is expressed as:
[0013] Pick up picture frames in the video data at a specified time interval, and identify the tracked target in the video data based on the set of picked-up picture frames;
[0014] ;
[0015] where: K(a) is the possibility that the dynamic target a in the video data is the tracked target; n is the total number of picked-up picture frames; sim(i, P take ) is the similarity between the image data of the dynamic target a in the i-th picture frame and the image data of the tracked target; λ is the correction; ω is the weight; m is the number of bins of the histograms of the dynamic target image and the tracked target image data in the picture frame; i(j), P take (j) are the normalized frequency values of the two histograms in the j-th bin; Δx, Δy are the distance differences in the horizontal and vertical directions between the dynamic target and the tracked target in their respective source picture frames; W, H are the width and height of the picture frame;
[0016] Among them, the time interval used to pick up picture frames in the video data is user-defined by the system end, and the maximum time interval does not exceed one second;
[0017] Based on the above formula, calculate the possibility that each dynamic target is the tracked target, and select the dynamic target pointed to by the maximum calculation result as the tracked target. The weight 1 > ω > 0, and the value of the weight ω is user-defined by the system end.
[0018] Furthermore, the perception layer includes an upload module, a perception module, and an extraction module. The upload module is used to upload the video data collected by the collection device during operation. The perception module is used to perceive the posture of the collection device, lock and track the target in combination with the video data in the upload module. The extraction module is used to receive the tracked target obtained by the operation of the perception module and extract the characteristic parameters of the tracked target;
[0019] Among them, the collection device is worn on the body surface by the user. The collection device is any electronic device with both video and audio recording functions and speed measurement functions. The perception module is integrated with a position sensor, and continuously operates based on the position sensor to perceive the moving direction, moving speed, and path of the collection device, that is, the posture of the collection device. The characteristic parameters of the tracked target extracted by the extraction module during operation include the tracked target image and the tracked target moving speed.
[0020] Furthermore, during the operation stages of the upload module and the perception module, after the perception module perceives the posture of the collection device, the upload module synchronously identifies the moving direction, moving path, and moving speed of each dynamic target in the video data. The perception module further compares the posture of the collection device with the moving direction, moving path, and moving speed of the dynamic target to lock the tracked target;
[0021] After the tracked target is locked, the extraction module synchronously extracts the image and moving speed of the tracked target from the video data;
[0022] The locking logic of the tracked target is expressed as:
[0023] ;
[0024] In the formula: f(a) is the probability that the dynamic target a in the video data is the tracked target; V a is the moving speed of the dynamic target a; V EQIPT is the moving speed of the collection device; γ is an adjustment factor; SIMM(L a ,L EQIPT ) is the similarity of the moving paths between the dynamic target a and the collection device; W a is the moving direction angle of the dynamic target a; W EQIPT is the moving direction angle of the collection device;
[0025] Among them, based on the above formula, the probability that each dynamic target in the video data is the tracked target is calculated, and the dynamic target corresponding to the maximum calculated value is selected as the tracked target. The adjustment factor γ takes a value of 1 or -1. When V a > V EQIPT , the adjustment factor γ = -1. When V a ≤ V EQIPT , the adjustment factor γ = 1.
[0026] Further, the similarity SIMM(L a ,L EQIPT ) between the dynamic target a and the moving path of the acquisition device is calculated by the following formula:
[0027] ;
[0028] In the formula: l a is the length of the moving path of the dynamic target a; S a is the bounding box area of the moving path of the dynamic target a; D a is the dispersion of the moving path of the dynamic target a; l EQIPT is the length of the moving path of the acquisition device; S EQIPT is the bounding box area of the moving path of the acquisition device; D EQIPT is the dispersion of the moving path of the acquisition device;
[0029] Among them, the above formula is used to calculate the similarity between the moving paths of each dynamic target and the acquisition device.
[0030] Further, after the control module runs to control the perception layer to lock the tracking target again, the feature parameters of the tracking target locked again are extracted synchronously, and compared with the feature parameters of the tracking target received by the receiving module. When the comparison result is consistent, the recognition module is triggered to run;
[0031] When the comparison result is inconsistent, the feature parameters of the tracking target received by the receiving module are iterated with the feature parameters of the tracking target locked again, and the perception layer is controlled to run to lock the tracking target again, and the feature parameters of the tracking target are extracted again. This continuous operation is performed until the feature parameters of the tracking target obtained by the operation of the control module are consistent with the feature parameters of the tracking target received by the receiving module, and then the operation ends and the recognition module is triggered to run.
[0032] Further, the value of the correction λ is: ;
[0033] In the formula: V a , V take are the average moving speed and the tracking target moving speed determined by the dynamic target a based on each frame of the picture; MAX(V a , V take ) is the maximum value in the parentheses;
[0034] When the feature parameters of the tracking target received by the receiving module are compared with the feature parameters of the tracking target obtained by the control module controlling the perception layer to run again, the comparison target is the tracking target image data in the feature parameters, and the comparison logic is the same as the calculation logic of sim(i, P take );
[0035] Among them, the system-side user defines the consistency determination threshold. When the comparison result is greater than or equal to the consistency determination threshold, it is determined that the tracking target feature parameters received by the receiving module are consistent with the tracking target feature parameters obtained by the control module controlling the re-operation of the sensing layer. Otherwise, they are inconsistent.
[0036] Furthermore, the indication layer includes an identification module, an indication module, and a marking module. The identification module is used to obtain the frame in the latest tracking target source video data in the recognition layer, capture the moving path of the tracking target based on the frame in the pre-position video of the video data, and synchronously obtain the current position of the acquisition device itself. The indication module is used to receive the moving path of the tracking target and the current position of the acquisition device itself in the identification module, and use the current position of the acquisition device itself as a reference to indicate the direction of the tracking target. The marking module is used to receive the video data collected by the operation of the acquisition device and perform frame-level continuous marking on the tracking target in the video data.
[0037] Among them, when the marking module performs frame-level continuous marking on the tracking target in the video data, each frame image in the video data is used as the marking target, and the marking operation on the tracking target is performed in the video data based on the square box with the user-defined color on the system side.
[0038] Furthermore, during the operation stage of the indication module, the current position of the acquisition device itself is a point, and the moving path of the tracking target is a polyline. The longest line segment in the moving path of the tracking target that is farthest from the current position of the acquisition device itself is taken as the indication line segment. The point of the current position of the acquisition device itself is connected to and extended from the nearest endpoint on the indication line segment, and the included angle formed by the extension line and the indication line is the direction of the tracking target indicated by the indication module.
[0039] The indication module is built-in with an audio module. After the indication module operates to obtain the included angle formed by the extension line and the indication line, it broadcasts the included angle. The broadcast format is: G1 is deflected by G2, T;
[0040] Among them, G1 and G2 represent directions, and T represents the degree of the included angle.
[0041] Furthermore, the receiving module is connected to an extraction module through wireless network interaction. The extraction module is connected to a sensing module and an upload module through wireless network interaction. The receiving module is connected to a control module and an identification module through wireless network interaction. The identification module is connected to the identification module through wireless network. The identification module is connected to an indication module and a marking module through wireless network interaction.
[0042] In a second aspect, an audio-visual multi-modal recognition method includes the following steps:
[0043] Step 1: Based on the acquisition device, collect video data in real time. Based on the perception device, sense the real-time posture of the acquisition device, lock the tracking target in the video data, and synchronously extract the characteristic parameters of the tracking target after the tracking target is locked.
[0044] Step 2: Obtain the characteristic parameters of the tracking target, perform the locking operation of the tracking target again, synchronously extract the characteristic parameters of the tracking target locked again, and compare whether the obtained characteristic parameters of the tracking target are consistent with the characteristic parameters extracted after the tracking target is locked again.
[0045] Step 31: If they are inconsistent, repeat the locking operation of the tracking target until the comparison result is consistent.
[0046] Step 32: If they are consistent, identify the tracking target in the video data in real time.
[0047] Step 4: Obtain the recognition result of the tracking target in the latest video data, give a direction indication to the tracking target based on the position of the acquisition device itself, and synchronously perform frame-level annotation on the tracking target in the video data.
[0048] Adopting the technical solution provided by the present invention, compared with the known public technology, it has the following beneficial effects:
[0049] The present invention provides an audio-visual multi-modal recognition method and system. During the operation of the system, the video data collected by the acquisition device is processed with the acquisition device as the main body, the tracking target is obtained in the video data, and the clear tracking target is finally determined with the specified repeated locking logic. Further, with the frame-level image processing and analysis technology, the tracking target is continuously tracked, so as to provide an indication for the system-side user, bring an auxiliary tracking effect for the system-side user in the tracking scenario of the tracking target, ensure the smoother development of the tracking task, and at the same time, with the continuous annotation of the tracking target in the video data, provide reference and guidance for the subsequent tracking task, effectively improving the functionality and use of the acquisition device (recorder) in the actual use scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 It is a schematic structural diagram of an audio-visual multi-modal recognition system;
[0052] Figure 2 It is a schematic flow diagram of an audio-visual multi-modal recognition method;
[0053] Figure 3 This is the schematic diagram of the tracking target indication logic in the present invention. Detailed implementation manners
[0054] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] The present invention will be further described below with reference to the embodiments.
[0056] Embodiment 1:
[0057] An audio-video multi-modal recognition system in this embodiment, as Figure 1 shown, includes: a perception layer, a recognition layer, and an indication layer;
[0058] The video data collected by the collection device during operation is uploaded to the perception layer in real time. The perception layer synchronously perceives the posture of the collection device, locks the tracking target based on the collected video and the posture of the collection device, and synchronously extracts the characteristic parameters of the tracking target. The recognition layer further receives the characteristic parameters of the tracking target extracted in the perception layer, synchronously controls the perception layer to operate again to lock the tracking target, continuously recognizes the tracking target in combination with the tracking target locked again and the characteristic parameters of the tracking target, and the indication layer receives in real time the recognition result of the tracking target operated by the recognition layer, and continuously indicates the tracking target based on the recognition result;
[0059] The perception layer includes an upload module, a perception module, and an extraction module. The upload module is used to upload the video data collected by the collection device during operation. The perception module is used to perceive the posture of the collection device, lock the tracking target in combination with the video data in the upload module, and the extraction module is used to receive the tracking target obtained by the operation of the perception module and extract the characteristic parameters of the tracking target;
[0060] Among them, the collection device is worn on the body surface by the user, and the collection device is any electronic device with both video and audio recording functions and speed measurement functions. The perception module is integrated by a position sensor, and continuously operates based on the position sensor to perceive the moving direction, moving speed, and path of the collection device, that is, the posture of the collection device. The characteristic parameters of the tracking target extracted by the operation of the extraction module include the tracking target image and the tracking target moving speed;
[0061] During the operation of the upload module and the perception module, after the perception module senses the attitude of the acquisition device, the upload module synchronously identifies the moving direction, moving path, and moving speed of each dynamic target in the video data. The perception module further compares based on the attitude of the acquisition device with the moving direction, moving path, and moving speed of the dynamic target to lock the tracking target;
[0062] After the tracking target is locked, the extraction module synchronously extracts the image and moving speed of the tracking target from the video data;
[0063] The locking logic of the tracking target is expressed as:
[0064] ;
[0065] In the formula: f(a) is the probability that the dynamic target a in the video data is the tracking target; V a is the moving speed of the dynamic target a; V EQIPT is the moving speed of the acquisition device; γ is the adjustment factor; SIMM(L a ,L EQIPT ) is the similarity of the moving paths between the dynamic target a and the acquisition device; W a is the moving direction angle of the dynamic target a; W EQIPT is the moving direction angle of the acquisition device;
[0066] Among them, based on the above formula, the probability that each dynamic target in the video data is the tracking target is calculated, and the dynamic target corresponding to the maximum calculated value is selected as the tracking target. The adjustment factor γ takes a value of 1 or -1. When V a > V EQIPT , the adjustment factor γ = -1. When V a ≤ V EQIPT , the adjustment factor γ = 1.
[0067] The locking logic of the tracking target is defined by the above logical formula to ensure that the system's perception layer and recognition layer stably lock the tracking target.
[0068] The calculation formula for the similarity SIMM(L a ,L EQIPT ) of the moving paths between the dynamic target a and the acquisition device is:
[0069] ;
[0070] In the formula: l a is the length of the moving path of the dynamic target a; S a is the bounding box area of the moving path of the dynamic target a; D a is the dispersion of the moving path of the dynamic target a; l EQIPT is the length of the moving path of the acquisition device; S EQIPTis the area of the bounding box of the moving path of the acquisition device; D EQIPT is the dispersion of the moving path of the acquisition device;
[0071] Among them, the above formula is used for calculating the similarity between each dynamic target and the moving path of the acquisition device;
[0072] The recognition layer includes a receiving module, a control module and an identification module. The receiving module is used to receive the tracking target feature parameters extracted by the operation of the sensing layer. The control module is used to control the sensing layer to operate again to relock the tracking target. The identification module is used to receive video data and continuously identify the tracking target in the video data;
[0073] The continuous recognition logic for the tracking target in the identification module is expressed as:
[0074] Pick up picture frames in the video data based on a specified time interval, and identify the tracking target in the video data with the set of picked-up picture frames;
[0075] ;
[0076] In the formula: K(a) is the possibility that the dynamic target a in the video data is the tracking target; n is the total number of picked-up picture frames; sim(i, P take ) is the similarity between the image data of the dynamic target a in the i-th picture frame and the image data of the tracking target; λ is the correction; ω is the weight; m is the number of bins of the histograms of the dynamic target image and the tracking target image data in the picture frame; i(j), P take (j) are the normalized frequency values of the two histograms in the j-th bin; Δx, Δy are the distance differences between the dynamic target and the tracking target in the horizontal and vertical directions in their respective source picture frames; W, H are the width and height of the picture frame;
[0077] Among them, the time interval used to pick up picture frames in the video data is user-defined by the system end-user, and the maximum time interval does not exceed one second;
[0078] Based on the above formula, calculate the possibility that each dynamic target is the tracking target, and select the dynamic target pointed to by the maximum calculation result as the tracking target. The weight 1>ω>0, and the value of the weight ω is user-defined by the system end-user;
[0079] The value of the correction λ is: ;
[0080] In the formula: V a , V take are the average moving speed and the tracking target moving speed of the dynamic target a determined based on each picture frame; MAX(V a , V take ) is the maximum value within the brackets;
[0081] When the tracking target feature parameters received by the receiving module are compared with the tracking target feature parameters obtained by the control module controlling the re - operation of the perception layer for consistency, the comparison target is the tracking target image data in the feature parameters, and the comparison logic is the same as the calculation logic of sim(i, P take )
[0082] Among them, the system - side user defines the consistency determination threshold. When the comparison result is greater than or equal to the consistency determination threshold, it is determined that the tracking target feature parameters received by the receiving module are consistent with the tracking target feature parameters obtained by the control module controlling the re - operation of the perception layer; otherwise, they are inconsistent;
[0083] Through the above logical formula, it is determined whether each dynamic target in the video data is a tracking target, thus providing further operation data support for the operation of the indication layer.
[0084] The indication layer includes an identification module, an indication module, and a marking module. The identification module is used to obtain the frame in the video data of the latest tracking target source in the recognition layer, capture the movement path of the tracking target based on the leading - position video in the video data of the frame, and synchronously obtain the current position of the acquisition device itself. The indication module is used to receive the movement path of the tracking target and the current position of the acquisition device itself in the identification module, and use the current position of the acquisition device itself as a reference to indicate the direction of the tracking target. The marking module is used to receive the video data collected by the operation of the acquisition device and perform frame - level continuous marking on the tracking target in the video data;
[0085] Among them, when the marking module performs frame - level continuous marking on the tracking target in the video data, each frame image in the video data is used as the marking target, and the marking operation on the tracking target is performed in the video data based on a square box with a user - defined color on the system side;
[0086] During the operation stage of the indication module, the current position of the acquisition device itself is a point, and the movement path of the tracking target is a polyline. The longest line segment in the movement path of the tracking target from the current position of the acquisition device itself is taken as the indication line segment. The current position point of the acquisition device is connected to and extended from the nearest endpoint on the indication line segment, and the included angle formed by the extension line and the indication line is the direction of the tracking target indicated by the indication module;
[0087] The indication module is built - in with an audio module. After the indication module operates to obtain the included angle formed by the extension line and the indication line, it broadcasts the included angle, and the broadcast format is: G1 deviates from G2, T;
[0088] Among them, G1 and G2 represent directions, and T represents the degree of the included angle;
[0089] The receiving module is connected with an extraction module through wireless network interaction. The extraction module is connected with a sensing module and an uploading module through wireless network interaction. The receiving module is connected with a control module and an identification module through wireless network interaction. The identification module is connected with a discrimination module through wireless network. The discrimination module is connected with an indication module and a marking module through wireless network interaction.
[0090] In this embodiment, the uploading module runs to upload the video data collected by the collection device. The sensing module synchronously senses the posture of the collection device, combines the video data in the uploading module to lock and track the target. The extraction module runs later to receive the tracked target obtained by the operation of the sensing module, and extracts the characteristic parameters of the tracked target. The receiving module further receives the characteristic parameters of the tracked target extracted by the operation of the sensing layer. The synchronous control module controls the sensing layer to run again to lock the tracked target again. Then the identification module receives the video data and continuously identifies the tracked target in the video data. The discrimination module runs in the identification layer to obtain the frame in the video data of the latest tracked target source. Based on the frame, it captures the moving path of the tracked target in the pre-position video of the video data, and synchronously obtains the current position of the collection device itself. The indication module synchronously receives the moving path of the tracked target and the current position of the collection device itself in the discrimination module, and uses the current position of the collection device itself as a reference to indicate the direction of the tracked target. Finally, the marking module receives the video data collected by the collection device and performs frame-level continuous marking on the tracked target in the video data;
[0091] Through the operation of the system in the above embodiment, with the collection device (recorder) as the main body, it assists the staff to get guidance and assistance in the scenario of tracking the tracked target, so as to ensure that the tracking task is carried out more effectively and avoid the situation of losing the tracked target as much as possible;
[0092] See Figure 3 As shown, based on the arrow indication in the figure, a logical display of the indication content run by the indication module is generated. From top to bottom based on the arrow indication in the figure, the point represents the current position of the collection device itself, and the multi-segment line composed of dotted lines represents the moving path of the tracked target. By intercepting the moving path of the tracked target, the point of the current position of the collection device itself is connected and extended with the intercepted line segment, and an included angle is obtained, which is also the source support for the indication content run by the indication module.
[0093] As Figure 1 As shown, after the control module runs to control the sensing layer to lock the tracked target again, it synchronously extracts the characteristic parameters of the re-locked tracked target and compares them with the characteristic parameters of the tracked target received by the receiving module. When the comparison result is consistent, it triggers the operation of the identification module;
[0094] When the comparison result is inconsistent, the characteristic parameters of the tracking target received by the characteristic parameter iterative receiving module of the re-locked tracking target are used to control the perception layer to run the locked tracking target again, and the characteristic parameters of the tracking target are extracted again. This continuous operation is carried out until the characteristic parameters of the tracking target obtained by the control module operation are consistent with the characteristic parameters of the tracking target received by the receiving module, and then the operation ends and the recognition module is triggered to run.
[0095] Through the above settings, further operational data support is provided for the operation of the recognition layer of the system in the above embodiment, ensuring that the linkage operation of each operation layer of the system in the above embodiment is more stable and continuously and effectively tracking the tracking target.
[0096] Embodiment 2:
[0097] At the specific implementation level, on the basis of Embodiment 1, this embodiment refers to Figure 2 to further specifically describe a multi-modal audio-visual recognition system in Embodiment 1:
[0098] A multi-modal audio-visual recognition method includes the following steps:
[0099] Step 1: Based on the acquisition device, video data is collected in real time. Based on the perception device, the real-time posture of the acquisition device is perceived. A tracking target is locked in the video data. Synchronously, after the tracking target is locked, the characteristic parameters of the tracking target are extracted.
[0100] Step 2: Obtain the characteristic parameters of the tracking target, perform the locking operation of the tracking target again, synchronously extract the characteristic parameters of the tracking target that is locked again, and compare whether the obtained characteristic parameters of the tracking target are consistent with the characteristic parameters extracted after the tracking target is locked again.
[0101] Step 31: If they are inconsistent, repeat the locking operation of the tracking target until the comparison result is consistent.
[0102] Step 32: If they are consistent, the tracking target is recognized in the video data in real time.
[0103] Step 4: Obtain the recognition result of the tracking target in the latest video data, give a direction indication to the tracking target based on the position of the acquisition device itself, and synchronously perform frame-level annotation on the tracking target in the video data.
[0104] In summary, during the operation of the system in the above embodiments, the video data collected by the acquisition device is processed with the acquisition device as the main body. The tracking target is obtained from the video data, and a clear tracking target is finally determined with the specified repeated locking logic. Further, with the frame-level image processing and analysis technology, the tracking target is continuously tracked, so as to provide an indication to the system-side user. In the tracking scenario of the tracking target, an auxiliary tracking effect is brought to the system-side user, ensuring that the tracking task is carried out more smoothly. At the same time, with the continuous annotation of the tracking target in the video data, reference and guidance are provided for subsequent tracking tasks, effectively improving the functionality and use of the acquisition device (recorder) in the actual use scenario.
[0105] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio and video multimodal recognition system, characterized in that: include: Perception layer, recognition layer, indication layer; The video data collected by the acquisition device is uploaded to the perception layer in real time. The perception layer synchronously perceives the posture of the acquisition device, locks the tracking target based on the collected video and the posture of the acquisition device, and synchronously extracts the tracking target feature parameters. The recognition layer further receives the tracking target feature parameters extracted from the perception layer, and synchronously controls the perception layer to run and lock the tracking target again. The tracking target is continuously identified by combining the re-locked tracking target and the tracking target feature parameters. The indication layer receives the recognition result of the tracking target run by the recognition layer in real time, and continuously indicates the tracking target based on the recognition result. The recognition layer includes a receiving module, a control module and a recognition module. The receiving module is used to receive the tracking target feature parameters extracted by the perception layer. The control module is used to control the perception layer to run again and re-lock the tracking target. The recognition module is used to receive video data and continuously recognize the tracking target in the video data. The continuous recognition logic for the tracking target in the recognition module is expressed as: Picking up picture frames in the video data based on a specified time interval, and identifying a tracking target in the video data with a set of picked up picture frames; ; Where: K(a) is the possibility that the dynamic target in the video data is the tracking target a; n is the total number of picked-up frames; sim(i,P take ) is the similarity between the image data of the dynamic target and the image data of the tracking target in the i-th picture frame; λ is the correction; ω is the weight; m is the number of bins in the histogram of the dynamic target image and the tracking target image data in the picture frame; i(j), P take (j) is the normalized frequency value of the two histograms in the jth bin; Δx and Δy are the distance differences between the dynamic target and the tracking target in the horizontal and vertical directions in their respective source frames; W and H are the frame width and height; The time interval used to pick up the picture frames in the video data is customized by the system end user, and the maximum time interval does not exceed one second; Based on the above formula, the possibility of each dynamic target being a tracking target is calculated, and the dynamic target with the maximum calculation result is selected as the tracking target, with a weight of 1>ω>0, and the value of weight ω is customized by the system user; After the control module runs to control the perception layer to lock the tracking target again, the feature parameters of the re-locked tracking target are extracted synchronously, and the feature parameters of the tracking target are compared with those received by the receiving module. When the comparison result is consistent, the recognition module is triggered to run; When the comparison result is inconsistent, the characteristic parameters of the tracking target received by the receiving module are iterated with the characteristic parameters of the tracking target locked again, and the perception layer is controlled to run and lock the tracking target again, and the characteristic parameters of the tracking target are extracted again. This operation is continued until the tracking target characteristic parameters obtained by the control module are consistent with the tracking target characteristic parameters received in the receiving module, and then the operation is terminated, triggering the recognition module to run; The modified λ value is: ; Where: V a 、V take is the average moving speed of the dynamic target a determined based on each frame, and the moving speed of the tracking target; MAX(V a ,V take ) is the maximum value without brackets; When the tracking target feature parameters received by the receiving module are compared with the tracking target feature parameters obtained by the control module and the perception layer, the comparison target is the tracking target image data in the feature parameters. The comparison logic is the same as sim(i,P take ) has the same logic; Among them, the system end user customizes the consistency judgment threshold. When the comparison result is greater than or equal to the consistency judgment threshold, it is determined that the tracking target feature parameters received by the receiving module are consistent with the tracking target feature parameters obtained by the control module controlling the perception layer to re-run. Otherwise, they are inconsistent.
2. The audio and video multimodal recognition system according to claim 1, characterized in that: The perception layer includes an upload module, a perception module and an extraction module. The upload module is used to upload the video data collected by the acquisition device. The perception module is used to perceive the posture of the acquisition device and lock the tracking target in combination with the video data in the upload module. The extraction module is used to receive the tracking target obtained by the perception module and extract the characteristic parameters of the tracking target. Among them, the acquisition device is worn on the body surface of the user. The acquisition device is any electronic device that has both video and audio recording functions and speed measurement functions. The perception module is integrated with a position sensor, and continuously runs based on the position sensor to perceive the moving direction, moving speed and path of the acquisition device, that is, the posture of the acquisition device. The extraction module runs to extract the tracking target feature parameters including the tracking target image and the tracking target moving speed.
3. The audio and video multimodal recognition system according to claim 2, characterized in that: During the operation phase of the upload module and the perception module, after the perception module perceives the posture of the acquisition device, the upload module synchronously identifies the moving direction, moving path and moving speed of each dynamic target in the video data, and the perception module further compares the posture of the acquisition device with the moving direction, moving path and moving speed of the dynamic target to lock the tracking target; After the tracking target is locked, the extraction module synchronously extracts the image and moving speed of the tracking target in the video data; The locking logic of the tracking target is expressed as: ; Where: f(a) is the probability that the dynamic target a in the video data is the tracking target; V a is the moving speed of the dynamic target a; V EQIPT is the moving speed of the acquisition device; γ is the adjustment factor; SIMM(L a ,L EQIPT ) is the similarity between the moving path of the dynamic target a and the acquisition device; W a is the moving direction angle of the dynamic target a; W EQIPT is the moving direction angle of the acquisition device; Among them, based on the above formula, the probability of each dynamic target in the video data being the tracking target is calculated, and the dynamic target corresponding to the maximum calculated value is selected as the tracking target. The adjustment factor γ takes the value of 1 or -1, and V a >V EQIPT When the adjustment factor γ=-1, V a ≤V EQIPT When , the adjustment factor γ=1.
4. The audio and video multimodal recognition system according to claim 3, characterized in that: The similarity SIMM (L a ,L EQIPT ) is calculated as: ; Where: l a is the length of the moving path of the dynamic target a; S a is the bounding box area of the moving path of the dynamic target a; D a is the discreteness of the moving path of the dynamic target a; l EQIPT S is the length of the moving path of the acquisition device; EQIPT D is the bounding box area of the moving path of the acquisition device; EQIPT The discreteness of the moving path of the acquisition device; The above formula is used to calculate the similarity of the moving paths of each dynamic target and the acquisition device.
5. The audio and video multimodal recognition system according to claim 1, characterized in that: The indication layer includes an identification module, an indication module, and an annotation module. The identification module is used to obtain the picture frame in the latest source video data of the tracking target in the identification layer, capture the moving path of the tracking target based on the front position video of the picture frame in the video data, and synchronously obtain the current position of the acquisition device. The indication module is used to receive the moving path of the tracking target and the current position of the acquisition device in the identification module, and use the current position of the acquisition device as a reference to indicate the direction of the tracking target. The annotation module is used to receive the video data collected by the acquisition device, and perform frame-level continuous annotation on the tracking target in the video data. Among them, when the labeling module continuously labels the tracking target at the frame level in the video data, each frame image in the video data is used as the labeling target, and the labeling operation is performed on the tracking target in the video data based on the box of the user-defined color on the system side.
6. The audio and video multimodal recognition system according to claim 5, characterized in that: During the operation phase of the indication module, the current position of the acquisition device is a point, and the moving path of the tracking target is a multi-segment line. The line segment farthest from the current position of the acquisition device in the moving path of the tracking target is taken as the indication line segment. The current position point of the acquisition device and the nearest end point on the indication line segment are connected and extended. The angle formed by the extension line and the indication line is the direction of the target tracking by the indication module. The indication module has a built-in audio module. After the indication module obtains the angle formed by the extended line and the indication line, the angle is broadcasted in the format of: G1 deviates from G2, T; Among them, G1 and G2 represent directions, and T represents the degree of the angle.
7. The audio and video multimodal recognition system according to claim 1, characterized in that: The receiving module is interactively connected to the extraction module via a wireless network, the extraction module is interactively connected to the perception module and the upload module via a wireless network, the receiving module is interactively connected to the control module and the recognition module via a wireless network, the recognition module is interactively connected to the identification module via a wireless network, and the identification module is interactively connected to the indication module and the labeling module via a wireless network.
8. A method for audio and video multimodal recognition, the method being an implementation method of an audio and video multimodal recognition system as claimed in any one of claims 1 to 7, characterized in that: The following steps are involved: Step 1: Collect video data in real time based on the acquisition device, perceive the real-time posture of the acquisition device based on the perception device, lock the tracking target in the video data, and extract the characteristic parameters of the tracking target after the tracking target is locked; Step 2: Obtain the characteristic parameters of the tracking target, perform the locking operation of the tracking target again, synchronously extract the characteristic parameters of the tracking target locked again, and compare whether the obtained characteristic parameters of the tracking target are consistent with the characteristic parameters extracted after locking the tracking target again; Step 31: If not consistent, repeat the locking operation of the tracking target until the comparison result is consistent; Step 32: Consistently, identify the tracking target in real time in the video data; Step 4: Obtain the latest target recognition results in the video data, indicate the direction of the target based on the position of the acquisition device, and simultaneously annotate the target at the frame level in the video data.
Citation Information
Patent Citations
Multi-modal personality characteristic identification method and system based on audio and video
CN118628958A
Comprehensive monitoring and management system for safe construction of intelligent seaport
CN118967063A
Method and apparatus for operating tracking algorithm, and electronic device and computer-readable storage medium
WO2022126415A1