Automatic wonderful event marking method based on sound source localization and visual fusion

By combining sound source localization and visual fusion technology with microphone arrays and cameras on mobile devices, the problem of traditional methods being unable to accurately locate exciting events in complex scenes has been solved, achieving high-precision and stable event localization and meeting the needs of real-time content generation.

CN121633994AActive Publication Date: 2026-03-10宁波工业互联网研究院有限公司 +1
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In complex scenarios, traditional visual and sound source localization methods struggle to accurately locate exciting events in environments with moving targets occlusion, changing lighting, and multiple sound sources. Especially in open spaces, existing technologies cannot accurately locate specific target areas in two-dimensional or three-dimensional space.

Method used

By combining a microphone array and a camera on a mobile device to acquire acoustic and visual data, acoustic event feature groups and spatial fusion feature groups are constructed. The array attitude matrix is ​​used to correct the sound source direction information, and the scene prior of visual information is combined to achieve the stability and accurate positioning of the sound source spatial direction.

Benefits of technology

It achieves multi-source fusion event localization in complex sports scenarios, breaks through the limitations of traditional methods, improves localization accuracy and stability, and can achieve high-confidence marking of exciting events at the regional or target level, meeting the needs of real-time content generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121633994A_ABST
    Figure CN121633994A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sound source localization, in particular to a wonderful event automatic marking method based on sound source localization and vision fusion. The method comprises the following steps: acquiring original acoustic data and original visual data, extracting an event voiceprint sequence in the original acoustic data, and determining scene prior information based on the original visual data; constructing an acoustic event feature group based on the original acoustic data and the event voiceprint sequence, and determining sound source direction information according to the microphone array; obtaining attitude data of the mobile terminal, and constructing an array attitude matrix; according to the invention, through multi-source fusion of acoustic positioning, mobile terminal attitude correction and visual space constraint, the stability of sound source space pointing in a dynamic shooting scene is effectively enhanced; the precision of wonderful event space positioning and the confidence degree of an automatic marking result are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound source positioning, in particular to a wonderful event automatic marking method based on sound source positioning and visual fusion. BACKGROUND

[0002] Early technologies mainly rely on visual detection, and infer the wonderful moment by identifying the abnormal motion, local acceleration or trajectory mutation of the moving object through image recognition. However, in complex scenes, the moving target is frequently blocked, the light changes significantly, and the camera field of view is limited, which makes it difficult for the visual method to meet the real-time content generation requirements. The mobile audio acquisition is seriously disturbed by the environment, and the background noise, reverberation and multipath reflection will cause the sound event positioning deviation. Especially in open spaces such as stadiums, multiple sound sources coexist and the reflection path is complex, and the traditional sound event recognition based on short-time energy or feature matching lacks spatial positioning capability and can only provide the time when the event occurs, but cannot determine the location where the event occurs.

[0003] At present, the industry gradually adopts multi-channel microphone array to carry out sound source direction estimation, and the typical method includes time difference of arrival estimation based on generalized cross-correlation (GCC). However, this kind of method still faces significant limitations when applied to mobile terminals: first, the posture of mobile devices changes constantly, and the array geometry continuously rotates relative to the world coordinate system, so the sound source direction cannot be kept stable; second, the scene structure has a significant impact on sound wave reflection, and the reflection peak interferes with the direct sound recognition, resulting in significant direction estimation error; third, relying only on the sound source direction is still insufficient to locate the specific target area in two-dimensional or three-dimensional space. SUMMARY

[0004] Therefore, it is necessary to provide a wonderful event automatic marking method based on sound source positioning and visual fusion to solve at least one of the above technical problems.

[0005] To achieve the above purpose, a wonderful event automatic marking method based on sound source positioning and visual fusion is applied to a mobile terminal, the mobile terminal includes a microphone array and a camera, and the method comprises the following steps: Step S1: obtaining original acoustic data and original visual data, extracting an event soundprint sequence in the original acoustic data, and determining scene prior information based on the original visual data; Step S2: constructing an acoustic event feature group based on the original acoustic data and the event soundprint sequence, and determining sound source direction information according to the microphone array; Step S3: obtaining posture data of the mobile terminal, and constructing an array posture matrix; correcting the sound source direction information by using the array posture matrix to generate sound source spatial pointing data; Step S4: Construct a spatial fusion feature group using the original visual data, scene prior information, and sound source spatial pointing data, and identify the spatial location of sound event triggering in the target scene area to form event triggering confidence. Step S5: Determine the results of highlight event marking based on the event trigger confidence.

[0006] This invention implements a multi-source fusion event localization mechanism in complex sports scenarios that is far superior to traditional single-mode audio and video recognition, with beneficial effects manifested on multiple levels. By constructing event acoustic signature sequences in the acoustic link and introducing prior scene information, it overcomes the limitations of early methods that relied solely on short-time energy or simple feature matching to locate the spatial position of sound sources, enabling sound events to have a traceable, structured representation. Through joint processing of multi-channel acoustic feature groups and microphone arrays, combined with a reflection path suppression mechanism, the sound source direction estimation remains stable even in multi-reflection environments. By mapping the sound source direction uniformly to the geodetic coordinate system using an array attitude matrix, changes in mobile device attitude no longer affect positioning accuracy, overcoming the problem of severe direction drift in traditional methods in mobile device scenarios. Furthermore, it introduces field information through visual information... The scene plane structure, boundary cage model, and moving target trajectory enable the spatial pointing data of sound sources to obtain external geometric constraints, no longer limited to the direction level, but able to achieve true spatial pointing at the region or target level; by binding sound trajectory with visual trajectory through spatiotemporal consistency measurement, misjudgment caused by visual occlusion, sudden changes in illumination, and multi-source mixing is significantly reduced; finally, by constructing spatial fusion feature groups, the temporal features, spatial features, and visual semantic features of sound are uniformly expressed, thereby achieving robust determination of the location of exciting events and outputting highly reliable labeling results to meet the needs of real-time content generation and automatic editing. Attached Figure Description

[0007] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Fig. 1 This is a flowchart illustrating the steps of an automatic event marking method based on sound source localization and visual fusion according to the present invention. Fig. 2 This is a schematic diagram of acoustic reflection path prediction based on visual priors; Fig. 3 This is a schematic diagram of sound source-visual spatial fusion and event localization. Detailed Implementation

[0008] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0009] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0010] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0011] To achieve the above objectives, please refer to Figs. 1 to 3 This invention provides an automatic event tagging method based on sound source localization and visual fusion, applicable to mobile devices. The mobile device includes a microphone array and a camera. The method includes the following steps: Step S1: Obtain raw acoustic data and raw visual data, extract event voiceprint sequences from the raw acoustic data, and determine scene prior information based on the raw visual data; Step S2: Construct acoustic event feature groups based on the original acoustic data and event voiceprint sequences, and determine the sound source direction information according to the microphone array; Step S3: Obtain the attitude data of the mobile device and construct the array attitude matrix; use the array attitude matrix to correct the sound source direction information and generate sound source spatial pointing data; Step S4: Construct a spatial fusion feature group using the original visual data, scene prior information, and sound source spatial pointing data, and identify the spatial location of sound event triggering in the target scene area to form event triggering confidence. Step S5: Determine the results of highlight event marking based on the event trigger confidence.

[0012] Preferably, before acquiring the raw acoustic data and raw visual data in step S1, the method further includes: Start the microphone array and camera on the mobile device, and initialize the acquisition of audio and video streams; Load a predefined acoustic event dictionary into the mobile device's memory; Initialize the inertial measurement unit (IMU) of the mobile device and establish the spatial calibration relationship between the IMU and the acoustic array.

[0013] In this embodiment of the invention, before step S1 is executed, the mobile terminal first starts the microphone array and camera and completes the initialization of the audio and video streams. The mobile terminal's built-in four-unit linear microphone array uses an 8mm unit pitch. Initialization commands are sent sequentially to the four microphone units through the driver interface. Each unit continuously collects sound pressure data for 200ms in a static environment and calculates the average level to form a noise reference. Then, the input gain is set according to the noise reference to keep each channel in a linear input range of -42dBV to -38dBV. The four channels are configured in synchronous acquisition mode with a sampling rate of 48kHz and a quantization accuracy of 16bit PCM. During audio stream initialization, a 4096-point circular buffer is created in memory for the four-channel data, and both the write and read pointers are set to 0. A hardware timer triggers a data write operation every 10.67ms, writing 512 audio blocks sequentially into the buffer and recording a monotonically increasing timestamp for each block for subsequent acoustic time difference measurements. When the camera is started, the mobile terminal... The wide-angle imaging unit enters the working state with a resolution of 1920×1080 and a frame rate of 60fps. It completes white balance locking (5000K) and exposure time fixing (1 / 240s) through the drive interface, and is configured with a dual-buffered image receiving area. The first buffer is used for the current frame acquisition, and the second buffer is used for subsequent visual processing. At the same time, the frame number and timestamp are registered after each frame acquisition is completed.

[0014] After the audio and video streams are prepared, the predefined acoustic event dictionary is loaded into the mobile device's memory. The acoustic event dictionary consists of sub-dictionaries for whistle events, ball-hitting sounds, and cheers. Each standard voiceprint entry contains 200ms long, 48kHz, 16-bit PCM format raw voiceprint data and a corresponding 512-dimensional frequency component array. The loading process reads the dictionary file sequentially in fixed 4KB blocks and performs a Fast Fourier Transform on each voiceprint segment in memory with a 512-point step size to obtain the standard frequency components of the event voiceprint. All voiceprint frequency components are written to contiguous memory in a matrix stacked manner, with each row corresponding to one standard voiceprint and each column corresponding to a fixed frequency component position. The three sub-dictionaries are arranged sequentially to form the comprehensive acoustic event dictionary.

[0015] After the acoustic event dictionary is loaded, the mobile device initializes the inertial measurement unit (IMU) and establishes the spatial calibration relationship between the IMU and the microphone array. The IMU includes a three-axis accelerometer and a three-axis gyroscope, with the accelerometer having a range of [missing information]. 4g with a resolution of 0.001g, gyroscope range of ±500° / s with a resolution of / s. The initialization process involves acquiring 500 frames of static output from the accelerometer and gyroscope respectively, calculating the three-axis zero-bias compensation values ​​and the three-axis angular velocity zero-drift compensation values, setting the update frequency of the inertial measurement unit to 200Hz, and simultaneously creating a data buffer queue of 200 frames in memory. During spatial calibration relationship construction, the center of the microphone array... Set the origin of the array coordinate system as the point where it is parallel to the origin of the inertial measurement unit. The fixed offset vector between them is set to (12mm, -4mm, 2mm). Based on the structural calibration data from the manufacturing stage, the IMU's... Axis and Array There are shafts The angle between the two, the IMU's Axis and Array There are shafts The angle between the two points is used to construct a rotation around the Z-axis. The rotation matrix Rz is then used to construct the rotation around the X-axis. The rotation matrix Rx is obtained, and the two are multiplied to obtain the orientation correction matrix. Then Combined with the above offset vector to form Homogeneous space transformation matrix .

[0016] Preferably, step S1 includes: Raw acoustic data and raw visual data are collected using the microphone array and camera on the mobile device, respectively. Event detection is performed on the raw acoustic data to identify audio segments containing potential sound events; Based on an acoustic event dictionary, feature matching and sparse coding are performed on audio segments, and event voiceprint sequences are extracted. Perform scene geometry analysis on the raw visual data to identify the main planar structures in the scene; Perform three-dimensional spatial orientation and parameter fitting on the main planar structure, and output the plane equation of the main planar structure in the geodetic coordinate system; Extract the boundary contours of the main planar structures and calculate the spatial scale to construct a three-dimensional geometric cage model describing the scene boundary; The planar equations and the three-dimensional geometric cage model are collectively defined as the scene prior information.

[0017] In this embodiment of the invention, after the microphone array and camera have completed initialization, the mobile terminal continuously acquires acoustic data at a 48kHz sampling rate and 16-bit PCM format using a four-unit linear microphone array, and then... Wide-angle camera with Visual images are acquired synchronously at a resolution and a frame rate of 60fps. Acoustic data is continuously written to a 4096-point circular buffer in units of 512 points, and the images output by the camera are stored in a double-buffered structure. Step S1 first extracts each 2048-point acoustic segment from the audio buffer and scans it within a time window. An amplitude threshold combined with short-term energy mean is used to identify acoustic energy surge intervals. The threshold is fixed at 3 times the static noise baseline, and the energy mean window length is 128 points. 2048-point segments that meet the threshold condition are marked as potential sound event segments. These segments are then matched with a predefined acoustic event dictionary. Each 2048-point data segment first undergoes a Fast Fourier Transform with a step size of 512 points to obtain its frequency components. Then, the corresponding dimension of the standard frequency component matrix of the voiceprint is read from the dictionary. The matching error value is obtained by calculating the absolute value of the difference point by point for each frequency component and summing them. The standard voiceprint with the smallest error value is determined to be the event type to which the segment belongs. The standard frequency components of all event segments are concatenated in their original time order to form an event voiceprint sequence.

[0018] After the voiceprint sequence is generated, step S1 performs scene geometry analysis on the visual data to construct the spatial constraints required for acoustic and visual fusion. Each frame of the visual image is... The grid division method generates fixed partitions. Gradient value statistics are performed in each partition, and regions with gradient direction distribution variance below a set threshold (threshold set to 0.15) are identified as potential planar candidate regions. The pixel coordinates of all candidate regions are then converted into a 3D point set in the camera coordinate system. The depth value is calculated using a fixed focal length model with a focal length of 4mm and a photosensitive element size of 1 / 2.3 inches. The obtained 3D point set is fitted with planar parameters using the least squares method. By constructing a planar equation of the form Ax + By + Cz + D = 0, the X, Y, and Z coordinates of the point set are substituted, and the planar parameters A, B, C, and D that satisfy the minimum sum of squared errors are obtained by multi-point summation and partial derivative calculation.

[0019] After determining the planar equations, boundary contours are formed by extracting boundary pixels belonging to the main planar structures. Boundary extraction is achieved by detecting gradient direction changes between meshes; when the difference in gradient variance between adjacent meshes is greater than 0.2, it is marked as a boundary segment. All segments are stitched together according to continuity rules to form a complete closed boundary. Next, the spatial scale of the 3D point set corresponding to the boundary segments is calculated. The length and width distribution of the boundary curve are obtained by calculating the Euclidean distance between adjacent points and accumulating it along the boundary direction. The 3D coordinates of the boundary point set are used to construct bounding boxes in the height and lateral directions, respectively. A 3D cage model is generated by selecting points on the surface of the bounding boxes and constructing connecting structures along the boundary direction. The cage model contains six faces and twelve edges, used to describe the spatial boundaries of the main planar structures. Finally, the above planar equations and the 3D cage model are combined to define the scene prior information.

[0020] Preferably, feature matching and sparse coding of audio segments based on an acoustic event dictionary includes: The system retrieves multiple event sub-dictionaries related to sports events from the acoustic event dictionary. These event sub-dictionaries include a whistle sound sub-dictionary, a ball-hitting sound sub-dictionary, and a cheering sound sub-dictionary. The audio segments are converted into time-frequency feature representations to form the input signal to be processed; The input signal is sparsely decomposed based on an acoustic event dictionary to extract sparse components representing different types of sound events. Based on the category attributes and energy distribution of sparse components, identify the types and corresponding intensities of exciting sports events contained in audio clips; Extracting event components related to exciting events in sports events from sparse components; The event voiceprint signal sequence is reconstructed based on the event sub-dictionary and event components.

[0021] In this embodiment of the invention, audio segments are captured by a microphone array in fixed-length units of 2048 points, with a sampling rate of 48kHz. To perform acoustic event recognition, three event sub-dictionaries related to sports events are first retrieved from system memory: a whistle sub-dictionary, a ball-hitting sound sub-dictionary, and a cheering sound sub-dictionary. Each sub-dictionary contains a standard voiceprint entry of 200ms in length, stored in a 512-dimensional frequency component structure. To process the audio segments, the 2048-point audio segments are first converted into time-frequency feature representations. Specifically, a Fast Fourier Transform is performed with a fixed window of 512 points, and adjacent windows share 256 points of data using a 50% overlap method, thereby generating four frequency components, each containing 512 frequency amplitude values. These four frequency components are concatenated in chronological order into a 2048-dimensional composite frequency vector.

[0022] To extract audio components associated with three event sub-dictionaries from the input signal, a sparse decomposition is performed using a line-by-line comparison approach. This operation does not use any model but directly compares the component differences between the input signal and the dictionary entries. First, for the standard voiceprint in each sub-dictionary, its 512-dimensional standard frequency components are repeated four times to align with the 2048-dimensional structure of the input signal. Then, the absolute difference is calculated for each frequency dimension and accumulated to obtain the total difference; the smaller the difference, the higher the matching degree. To form sparse components, only the three voiceprint entries with the smallest differences in each sub-dictionary are retained, and the standard frequency components of these entries are weighted according to the reciprocal of the error to calculate a weighted average frequency component, which serves as the sparse component for the corresponding event type. This allows for the extraction of a 2048-dimensional weighted frequency vector for each event type, representing the whistle component, the ball-hitting sound component, and the cheering sound component, respectively.

[0023] Once the sparse components are identified, the types and intensities of sporting events contained in the audio clips are determined based on the category attributes of these components and their energy distribution. To measure the energy distribution, the sum of the squares of the 2048-dimensional frequency indices of each sparse component is used as the energy value, and the relative proportions between the three energy values ​​are compared. If the energy proportion of a certain event type exceeds 40% of the total energy, that event type is determined to be the dominant sound event; if the energy proportions of two event types are both higher than 20%, a composite event is identified in the clip. Event intensity is classified in a fixed stratification based on the absolute magnitude of the sparse component energy values, where energy values ​​greater than 2000 are classified as high intensity, 1000 to 2000 as medium intensity, and less than 1000 as low intensity.

[0024] After identifying the event type and intensity, frequency content directly related to the event is further extracted from the sparse components. The extraction process involves filtering high-amplitude frequency points in the top 20% of the 2048-dimensional weighted frequency vector to obtain a set of key frequency components, which are considered the core frequency features representing the event type. Subsequently, based on the original time series structure of similar events in the event sub-dictionary, the key frequency components are redistributed back to four time windows, corresponding to four consecutive frequency frames. These four frequency frames are arranged in chronological order and restored to time-domain waveforms using inverse fast Fourier transform, thereby reconstructing the event audioprint signal sequence.

[0025] Preferably, step S2 includes: The temporal energy envelope and frequency spectral features of sound events are calculated based on the event voiceprint sequence to form an initial event feature vector; Multi-channel joint analysis of raw acoustic data is performed to determine the time difference information of sound waves arriving at different microphones; The candidate set for analyzing the direction of the sound source is based on time difference information. The candidate set contains multiple direction candidates formed by direct sound and reflected sound. The initial event feature vector and the candidate set of sound source directions are fused to form an acoustic event feature group; The sound source direction information is determined based on the confidence weight of each candidate direction in the acoustic event feature group.

[0026] In this embodiment of the invention, after the event voiceprint sequence is generated, the time-domain energy envelope and frequency-domain spectral features of the sound event are first calculated based on the sequence. The event voiceprint sequence contains four consecutive time-frequency recovered time-domain waveforms, each with a length of 512 points and a sampling rate of 48kHz. The energy envelope is obtained by summing the absolute values ​​of each waveform segment using a moving average, with a moving window length of 32 points and a step size of 16 points. 50% of the data is retained between adjacent windows, ensuring that the envelope curve reflects the change in event intensity over time. The frequency-domain spectral features are obtained by performing a 512-point Fast Fourier Transform on each waveform segment to obtain a 512-dimensional frequency amplitude array. The four frequency arrays are then sequentially concatenated to form a 2048-dimensional frequency vector. The statistical values ​​(maximum, mean, and variance) of the energy envelope are concatenated with the frequency vector to form an initial event feature vector of length 2051.

[0027] After constructing the initial event feature vector, multi-channel joint analysis was performed on the raw acoustic data to determine the time difference of sound waves arriving at different microphones. The original four-channel audio was synchronously acquired at a sampling rate of 48kHz, and each 4096-point audio segment was used for time difference extraction. First, the data of the four channels were normalized to unify the amplitude range to [-1, 1]. Then, cross-correlation was performed on any two channels (such as channel 1 and channel 2). The signals of channel 1 and channel 2 were successively delayed and multiplied point by point and summed, with the delay amount scanning from -128 points to 128 points. When the cross-correlation value under a certain delay reaches its maximum value, this delay amount is taken as the arrival time difference of channel 1 relative to channel 2. A total of six sets of time difference measurements were formed between the four channels, with each set measuring within the range of [-128, 128] points, corresponding to a physical delay interval of [-2.67ms, 2.67ms].

[0028] Based on the aforementioned time difference information, a candidate set of sound source directions is analyzed. The four-unit linear microphone array is geometrically positioned in a straight line on the device body, with a unit spacing of 8mm. Assuming sound waves propagate at a speed of 343m / s, the sound path difference can be obtained by multiplying the time difference between each microphone pair by 343m / s. For a linear array, the sound path difference d can be calculated using the formula... Calculate the horizontal angle of incidence corresponding to the direct sound. Because of reflected sound in sports venues, the time difference may include reflected path signals in addition to the direct sound. To construct a candidate set, the path difference offset caused by reflected sound needs to be included in the calculation while estimating the angle of the direct sound. The path difference of reflected sound is larger than that of direct sound, and its delay is usually distributed in the middle to high delay range within the above-mentioned [-128, 128] points. Therefore, in each set of time difference measurements, values ​​with absolute delay values ​​between 80 and 128 are retained, and the corresponding angles are converted using the same path difference formula to form candidate reflected sound directions. The direct sound candidates and reflected sound candidates are recorded in the form of angle lists, forming a candidate set for the sound source direction.

[0029] After obtaining the candidate set, the initial event feature vector is fused with the candidate set to form an acoustic event feature group. To achieve the fusion operation, for each candidate direction angle, a 2048-dimensional frequency vector in the frequency domain is extracted from the 2051-dimensional initial event feature vector, and weights are established for the candidate directions based on the distribution of frequency components. The weights are calculated by dividing the 2048-dimensional frequency vector into segments of 256 dimensions, calculating the energy value of each segment, mapping the energy value proportionally to the range of 0 to 1, and then constructing mapping coefficients based on the cosine relationship between the segment number and the candidate direction angle, so that the frequency energy contributes more to the frequency bands related to the direction of sound wave propagation. The above weights and the angle values ​​of the candidate directions are stored in the feature group in a one-to-one correspondence, finally forming an acoustic event feature group composed of angle-weight pairs.

[0030] After constructing the acoustic event feature set, the final sound source direction information is determined based on the confidence weights of each candidate direction within the feature set. The confidence score is obtained by summing the weights of the energy segments corresponding to each candidate direction. For direct sound candidates, the sum of the weights is directly used as their confidence score; for reflected sound candidates, the sum of the weights is multiplied by 0.6 as their confidence score attenuation factor to reduce the influence of reflected sound. The angle with the highest confidence score is output as the final sound source direction information.

[0031] Of particular importance is that, after extracting the peak sequence from the generalized cross-correlation function, the following is also included: Calculate the first-order reflection path delay of sound waves reflected from the scene boundary based on the scene prior information, and construct a set of possible reflection delays; Compare the time delay corresponding to each peak in the peak sequence with the set of suspected reflection time delays, and output the matching degree comparison results. Based on the matching degree comparison results, suppression weights are applied to the peaks in the peak sequence that belong to the reflection path, generating candidate arrival time difference values ​​after reflection suppression correction.

[0032] In this embodiment of the invention, the cross-correlation curves of each microphone pair are scanned through a range of [-128, 128] points to form delay-correlation value pairs. The first five peaks are extracted after arranging the correlation values ​​from largest to smallest, forming a peak sequence. Each peak is recorded with its corresponding delay point position (unit: sampling point). To identify the reflection path components, the first-order reflection path delay of the sound wave on the scene boundary is first calculated based on the scene prior information constructed in step S1. The scene prior information includes the plane equations of the main planar structures in the geodetic coordinate system. And a 3D geometric cage model describing the scene boundaries. Centered on the microphone array... Starting from the direction of the direct sound candidate Construct a ray with direction vector. Where u is the unit direction vector and t is a positive real scalar. The ray intersects with each facet of the cage-like model piece by piece. The intersection method involves substituting the L(t) coordinates into the explicit equation of the plane containing the facet to obtain the intersection point parameter t_i, and then determining whether the intersection point lies within the projection range of the facet boundary. The smallest value among all valid t_i corresponds to the first contact point of the sound wave from the array center to the boundary, denoted as P_ref. The sound wave undergoes specular reflection at point P_ref, and the reflection direction is symmetrically operated on by the vector u along the boundary plane normal vector n, i.e. The second segment of the reflection path originates from P_ref along u_ref and continues until it reaches the target channel position O_k of the microphone array again. The length of the second segment is obtained through analytical geometry. The theoretical time delay of the first-order reflection path is obtained by summing the lengths of the two segments and dividing by the speed of sound, 343 m / s. The value of _ref (in seconds) is multiplied by the sampling rate of 48kHz to convert it into the number of sampling points, forming the suspected reflection delay set. Since there are multiple combination paths between the reflection direction and the channel position, the suspected reflection delay set usually contains 2 to 4 discrete delay points, each of which is recorded as an integer sampling point.

[0033] After constructing the set of suspected reflection delays, the matching degree of each peak delay point in the generalized cross-correlation peak sequence is compared with each suspected delay in the set. The matching degree is calculated by taking the absolute difference between each peak delay point d_p and each suspected delay d_r. =|d_p−d_r|, if If there are 5 or fewer sampling points, the peak value is considered to have a high degree of consistency with a suspected reflection path, and the matching degree value M = 1− / 5 is recorded as the matching degree of this peak; if If the number of sampling points is greater than 5, the matching degree is recorded as 0. The maximum value of the matching degree of each peak is taken to form the final matching degree comparison result M_p of that peak, and the range of M_p is [0,1]. When M_p is close to 1, it means that the peak height is close to the reflection time delay; when M_p is close to 0, it means that the peak does not have reflection path characteristics and is directly at the sound peak.

[0034] Based on the matching degree comparison results, reflection path suppression is performed on the peak sequence. The suppression method involves multiplying the peak correlation value by an attenuation weight, defined as W = 1 − 0.7·M_p. That is, when the peak perfectly matches the reflection path (M_p = 1), its correlation value is multiplied by 0.3; when the peak does not match, the original value is retained (M_p = 0). The suppressed correlation values ​​are then reordered, and the peak with the highest correlation value is extracted as the candidate value for the arrival time difference of the direct sound. For cases requiring multiple candidate values ​​(e.g., when the sound field is complex and the top three peaks need to be retained), the same principle is used to extract the suppressed peak sequence, so that all reflection delays are stably suppressed, thereby generating the arrival time difference candidate value corrected by reflection suppression.

[0035] Preferably, step S3, which involves acquiring the mobile device's pose data and constructing the array pose matrix, includes: The inertial measurement unit of the mobile device acquires the three-axis acceleration and three-axis angular velocity data of the mobile device in the geodetic coordinate system in real time. The tilt angle of the mobile end relative to the direction of gravity is calculated based on triaxial acceleration data to determine the pitch angle and roll angle; The dynamic offset is compensated based on the three-cycle angular velocity data, the azimuth angle is corrected, and the yaw angle is output. Calculate the rotation matrices around the Z-axis, Y-axis, and X-axis based on the pitch angle, roll angle, and yaw angle respectively, and generate the complete rotation transformation matrix; The complete rotation transformation matrix is ​​determined as the array attitude matrix for sound source spatial pointing correction.

[0036] In this embodiment of the invention, the mobile device's built-in inertial measurement unit simultaneously reads three sets of acceleration data (a_x, a_y, a_z) from the triaxial accelerometer and three sets of angular velocity data (g_x, g_y, g_z) from the triaxial gyroscope at a sampling frequency of 200Hz. The triaxial acceleration data undergoes a single-pass mean filtering process, where the mean filtering window length is limited to 3 sampling points. The filtered acceleration data is obtained by calculating the arithmetic mean of the sampled values ​​within the window, and is used to extract the direction of gravity. The tilt angle and pitch angle of the mobile device relative to the direction of gravity are calculated using the filtered acceleration data. Roll angle The yaw angle was calculated using trigonometric functions, with division and square root operations performed in double-precision floating-point mode, and the angle result stored in radians. Subsequently, dynamic offset compensation was performed on the yaw angle using three-axis angular velocity data acquired by the gyroscope, with the compensation method implemented in each IMU sampling cycle. Within t=1 / 200s, set the yaw angle increment to g_z t, and add it to the accumulated yaw angle of the previous sampling period to form the real-time yaw angle. To compensate for the instantaneous drift caused by rapid shaking of the mobile device, a first-order attenuation coefficient k=0.92 is added to the yaw angle increment, through... _new= _old k+g_z Azimuth correction is performed using the 't' method, ensuring the output yaw angle remains continuous and trackable. At the pitch angle... Roll angle With yaw angle After the calculation is completed, rotation matrices around the Z-axis, Y-axis, and X-axis are constructed according to the definition of the spatial coordinate system. The rotation matrix R_z around the Z-axis is defined by [[cos , −sin , 0], [sin cos Constructed using [0, 0, 1]; the rotation matrix R_y around the Y-axis is defined by [[cos 0], ... , 0,sin ], [0, 1, 0], [−sin , 0, cos Construct; rotate the matrix R_x around the X-axis with [[1, 0, 0], [0, cos...] , −sin ], [0, sin cos Construct. The three rotation matrices are arranged according to... The matrix multiplications are performed sequentially, with all matrix multiplications executed in double-precision floating-point mode, ultimately yielding... Complete rotation transformation matrix. This rotation transformation matrix is ​​recorded as the array attitude matrix.

[0037] Preferably, step S3, which involves correcting the sound source direction information using the array attitude matrix, includes: Obtain the static geometric layout coordinates of the microphone array in the mobile device's body coordinate system; The static geometric layout coordinates are spatially transformed using the array attitude matrix to generate the dynamic spatial coordinates of the microphone array in the geodetic coordinate system. Based on dynamic spatial coordinates, coordinate system transformation is performed on the sound source direction information to generate sound source spatial pointing data that is independent of the mobile device's posture.

[0038] In this embodiment of the invention, the static geometric layout coordinates of the microphone array in the mobile device's body coordinate system are first read. The static geometric layout is given by the mobile device's factory calibration data, with the origin O_d of the body coordinate system as the reference point. The center point position of each microphone is represented using three-dimensional coordinates, denoted as . and All units are meters. The four coordinate points are arranged in a rectangular array, with the spacing between adjacent microphones limited to 0.035m, and the arrangement direction is strictly aligned with the X and Y axes of the camera's coordinate system. Then, using... The array attitude matrix R performs a spatial transformation on each static coordinate point. Specifically, it represents each point as a three-dimensional column vector and performs matrix multiplication. The multiplication uses double-precision floating-point calculations to ensure that each microphone obtains a unique dynamic spatial coordinate in the geodetic coordinate system. The above transformation is equivalent to rotating the entire microphone array in three dimensions around the origin of the device's coordinate system, aligning its orientation with the current posture of the mobile device. After completing the spatial transformation of all microphones, the sound source direction information obtained in step S2 (using the incident direction vector in the microphone array coordinate system) is used. (Represented), a coordinate system transformation consistent with the spatial coordinates is performed on it to generate the attitude-corrected spatial pointing vector of the sound source. Since vector multiplication only reflects directional relationships, after rotation... Perform unitization, i.e., divide by its modulus || ‖, thus obtaining the final sound source spatial pointing data d_norm= / ‖ This vector is located in the geodetic coordinate system and is independent of the mobile device's attitude.

[0039] Preferably, the spatial transformation of the static geometric layout coordinates using the array attitude matrix includes: Perform matrix multiplication between the static geometric layout coordinates and the array attitude matrix to calculate the preliminary spatial coordinates of the microphone array in the geodetic coordinate system at the current moment; Based on real-time displacement data from the mobile device, the initial spatial coordinates are translated and compensated to output the dynamic spatial coordinates of the microphone array.

[0040] In this embodiment of the invention, the static geometric layout coordinates of the microphone array in the mobile device's body coordinate system are organized point by point into a three-dimensional column vector, denoted as . , , , The four coordinate points are provided by the factory calibration and form a rectangular array around the origin of the fuselage coordinate system. The horizontal spacing is set to 0.035m and the vertical spacing is set to 0.020m. All coordinates are in meters. To map the array attitude to the geodetic coordinate system, a matrix multiplication operation is performed between each static coordinate point and the array attitude matrix. The matrix multiplication employs double-precision floating-point calculation mode, and the multiplication order strictly maintains the array attitude matrix on the left and the coordinate point on the right, ensuring that the rotation operation unfolds around the fuselage coordinate origin. The matrix multiplication output... This represents the initial spatial coordinates of the array after rotating around the center of the fuselage at the current moment. Subsequently, real-time displacement data provided by the mobile device's motion sensors is used to compensate for the translation of these initial coordinates. The real-time displacement data is obtained by joint integration from the accelerometer, gyroscope, and magnetometer, and then uniformly converted into a three-dimensional displacement vector in the geodetic coordinate system. The unit is meters, and the integration time step is 5 ms. Translation compensation is achieved by adding this displacement vector to the initial spatial coordinates, i.e. In this way, each microphone point is transformed from its internal coordinates to dynamic spatial coordinates in the geodetic coordinate system, and its spatial position is updated in real time as the mobile device moves in the sports scene.

[0041] Preferably, step S4 includes the following steps: Perform image recognition on raw visual data to detect and locate one or more potential event-triggered areas in the target scene; The spatial pointing data of the sound source is back-projected onto the image coordinate system composed of the original visual data to generate the sound source pointing projection line; Identify the specific sound event type corresponding to the event voiceprint sequence; Based on prior scene information, determine whether the spatial location pointed to by the projection line of the sound source matches the semantic type of the potential event triggering area; Within a continuous time period, the dynamic changes in the spatial pointing data of the sound source are tracked to generate the sound source pointing trajectory. Calculate the spatiotemporal consistency measure between the sound source pointing trajectory and the visual motion trajectory of the moving target within the potential event triggering area; By combining spatial relationship metrics, visual features of potential event triggering areas, and stability indicators of sound source spatial pointing data, a spatial fusion feature set is constructed. The spatial fusion feature set is input into a pre-trained event localization and discrimination model to calculate the probability that each potential event triggering region is the location of the real sound event. The target scene area and the corresponding event trigger confidence level are determined based on probability.

[0042] In this embodiment of the invention, image recognition is first performed on the raw visual data to detect potential event-triggered areas in the target scene. The raw visual data is acquired by a mobile camera at 60fps, with each frame containing an image of size [size missing]. The image recognition process employs a fixed-structure target detector, whose detection categories include athlete regions, ball game regions, referee action regions, and spectator seating regions. The detector outputs the two-dimensional bounding box coordinates [x_min, y_min, x_max, y_max] for each region and provides the semantic type of the region. By performing frame-by-frame detection on each frame of the image, a continuous sequence of potential event-triggered regions is obtained.

[0043] Then point the sound source space to the data. A back projection is performed to place it into the image plane coordinate system. The back projection uses the intrinsic parameter matrix K (focal lengths f_x, f_y, principal point coordinates c_x, c_y) and the extrinsic parameter matrix [E] (rotation and translation) obtained in step S1 calibration of the camera. The sound source direction vector is combined with the camera projection model to construct a three-dimensional ray. Where d_cam is the camera coordinate system direction vector obtained by rotating the sound source direction vector through the extrinsic parameter matrix, and O_c is the coordinate of the camera center point. Substituting r(t) into the projection equation... The two-dimensional image plane trajectory can then be obtained. The sound source pointing to the projection line is drawn along this trajectory, and a two-dimensional line segment with a resolution of 1 pixel is formed by sampling integer pixels.

[0044] After obtaining the projection lines, the sound event type is identified based on the event voiceprint sequence. After sparse component classification, the event voiceprint sequence corresponds to one of three categories: whistle event, ball-hitting sound event, or cheering event, and is accompanied by an instantaneous energy intensity parameter E_s, which is the normalized value of the sparse component energy in the range of 0–1, used for subsequent fusion determination.

[0045] Next, based on the scene's prior information, a consistency judgment is made between the location pointed to by the sound source projection line and the semantic type of the potential event triggering area. The scene's prior information includes the main planar structure, scene layout boundaries, and a 3D geometric cage model. The spatial positions of the projection line's extension direction are intersected by the cage structure. The intersection point is obtained by simultaneously solving the plane equations of the ray and the cage-like surface, yielding the landing point P_int of the projection line in 3D space. The spatial category of P_int is compared with the semantic type of the potential event triggering area. For example, a ball-hitting sound event requires pointing to the court area, while a whistle event requires pointing to the referee area. If the spatial category and semantic type are consistent, it is recorded as "Spatial Consistency Flag = 1"; otherwise, it is recorded as 0.

[0046] Subsequently, the dynamic changes in the spatial pointing data of the sound source were tracked over a continuous time period. The sound source direction vector was updated at 50Hz, and the direction vectors of 20 consecutive frames (400ms time window) were concatenated to form the sound source pointing trajectory, which consists of a series of unit vector point sets, each point denoted as d_norm(t_i). The amplitude of the angular change of the trajectory was detected. Record the temporal stability of the sound source direction.

[0047] Next, the trajectory of the moving target within the potential event triggering area is calculated on the visual side. Optical flow tracking is performed on the athlete or sphere in this area, selecting 150 uniformly distributed feature points within the area, and obtaining the two-dimensional motion trajectory frame by frame using Lucas-Kanade tracking. Using the three-dimensional calibration information of the camera, the two-dimensional trajectory is back-projected into three-dimensional space to obtain the visual motion trajectory v(t_i).

[0048] To measure the spatiotemporal consistency between the sound source trajectory and the visual trajectory, the directional angular consistency C_dir and the dynamic positional matching degree C_pos within the time window are calculated. C_dir is defined as the cosine of the average angular consistency over 20 frames, i.e. ,in This represents the angle between the sound source direction and the visual direction; C_pos is defined as the normalized form of the average Euclidean distance between the two trajectories under the frame correspondence relationship, i.e. Where P_src(t_i) is the sound source trajectory point, and P_vis(t_i) is the visual trajectory point. Together, these two constitute the spatiotemporal consistency metric C_sp = 0.5·C_dir + 0.5·C_pos.

[0049] Subsequently, the spatial relationship metric C_sp, the visual features of the potential event triggering region (including region area, region center location, number of visible edges, and feature point density), and the stability indicators of the sound source spatial pointing data (including the standard deviation of trajectory angle change) were used. and the speed of change of direction Combining these elements to construct spatial fusion feature groups .

[0050] The spatially fused feature set F is input into a pre-trained event localization and discrimination model. This model has a fixed three-layer fully connected structure, with the feature vector F as input and three probability values ​​as output. , , , representing the probabilities that the potential event triggering region is a ball-hitting sound event, a whistle event, and a cheering event, respectively. Based on the sound event type, the corresponding probability is read as the event trigger confidence level. If the current sound event category is a ball-hitting sound, then output . As a confidence level.

[0051] Finally, the area with the highest confidence level is selected from all potential event triggering areas as the target scene event location, and its confidence level is recorded as the basis for automatic marking of exciting events.

[0052] Of particular importance is the detection and location of one or more potential event-triggered areas in the target scene, including: Continuously track one or more moving targets within the potential event triggering area and generate the motion trajectory of each moving target; The spatial pointing data of the sound source is projected in reverse onto the image coordinate system to form a dynamic sound source pointing line; Calculate the spatiotemporal overlap between the pointing line of the dynamic sound source and the trajectory of each moving target within a continuous time window to generate a sound source-target correlation sequence. Based on the sound source-target correlation sequence, determine whether the sound source direction has a continuous matching relationship with a specific motion trajectory; When the determination result indicates the existence of the continuous matching relationship, the event trigger confidence of the potential event triggering area corresponding to the motion trajectory is increased.

[0053] In this embodiment of the invention, to detect and locate potential event-triggered areas in the target scene, all detected moving targets in the original visual data are first continuously tracked. The original visual data is a 60fps video stream, with each frame having a resolution of [resolution missing]. The moving targets include athletes, the ball, and the referee's action area. To obtain continuous motion trajectories, 120 feature points are uniformly sampled within the 2D bounding box of each moving target, and frame-by-frame displacement estimation is performed using optical flow tracing. The optical flow tracing process uses a double pyramid structure to calculate the pixel displacement of the feature points from frame t to frame t+1, and the mean coordinates of all feature points form the centroid position of the moving target in the current frame. By temporally sorting the centroid positions of consecutive frames, a motion trajectory sequence for each moving target can be constructed. , where k represents the frame number.

[0054] Then point the sound source space to the data. Back-projecting to the image coordinate system yields the sound source pointing line. Using the camera's intrinsic matrix K and extrinsic matrix [E], the back-projection transforms the sound source direction vector to the camera coordinate system via a rotation matrix, obtaining the direction vector d_cam. A ray is constructed starting from the camera center O_c. Substitute the ray into the projection equation , with step size At t=0.05m, point-by-point sampling is performed along the direction vector to generate a discrete pixel line segment composed of multiple pixel coordinates, called the sound source pointing line L_src.

[0055] To measure the spatial coupling relationship between the sound source direction and each motion trajectory, the spatiotemporal coincidence degree between the dynamic sound source pointing line and the motion trajectory of each moving target is calculated within a continuous time window. The time window is set to 400ms, corresponding to 24 frames. For the i-th frame in each window, the nearest distance d_min(i) between the sound source pointing line L_src(i) and the centroid point (x_i, y_i) of the moving target is calculated. The distance value is then mapped to the coincidence degree. ,in A 15-pixel value is used to map the distance difference to a continuous weight between 0 and 1. The average overlap is then calculated over the entire time window. , where N=24. A source-target correlation sequence is independently calculated for each moving target.

[0056] Subsequently, based on the correlation sequence, it is determined whether the sound source directionality has a continuous matching relationship with a specific motion trajectory. The determination method is to compare C_avg with a set threshold. _c, where _c=0.62. When C_avg When C_avg, it is assumed that the direction of the sound source is stably pointing to the trajectory of the moving target within this time window, forming a spatially continuous matching relationship, and the matching flag flag=1 is recorded; if C_avg If _c, then record flag=0.

[0057] If a persistent matching relationship is determined, the event trigger confidence of the potential event triggering region corresponding to the motion trajectory is increased. The confidence increase is achieved by updating the current confidence value p of that region to [value missing]. Where w is set to 0.35, it is used to map the spatiotemporal correlation strength to the confidence increment. After the confidence is updated, the region will be more likely to be identified as the real event trigger source in the subsequent event localization in step S4.

[0058] Preferably, step S5 includes the following steps: Compare the event trigger confidence level with a preset confidence threshold; When the event trigger confidence is greater than or equal to the confidence threshold, a valid exciting event is determined to have occurred. Obtain the target scene area and timestamp information corresponding to the relevant exciting events; Based on the target scene area and timestamp information, generate and output the results of highlight event marking.

[0059] In this embodiment of the invention, to determine whether a remarkable event is valid, the event trigger confidence level p_evt output from the preceding steps is first compared with a preset confidence threshold. The event trigger confidence level p_evt is compared with _evt. The event trigger confidence level p_evt is a floating-point number between 0 and 1, where a higher value indicates a higher confidence level in the event's occurrence. In this embodiment, the confidence threshold is... `_evt` is set to 0.72. This value is determined by the system based on long-term statistical results for sports events, and is used to ensure the stability and reliability of event triggering. When `p_evt` is greater than or equal to... When _evt is called, it is determined whether there are valid highlight events within the current time window; otherwise, the marking process is not entered.

[0060] When an event is determined to be a valid highlight, the target scene region and precise timestamp information to which the event belongs are read from the corresponding region record table in step S4. The target scene region is represented by a two-dimensional bounding box, in the format B=[x_min, y_min, x_max, y_max], with units of pixel coordinates. The timestamp information is provided by the video acquisition system, in milliseconds, and aligned with the audio acquisition timeline of the mobile device. In this embodiment, UTC time is used, and it is recorded as t_evt. The scene region information and timestamp are stored in an internal event buffer. Each update overwrites low-confidence event records to ensure that the output event is the valid event with the highest confidence within the current acquisition period.

[0061] Subsequently, event labeling results are generated based on the target scene region and timestamp. The event labeling results include four items: event type label_evt, event occurrence time t_evt, event location B, and event confidence p_evt. The labeling generation process is implemented by constructing structured event record units, with the following format: The event type `label_evt` is obtained from the event voiceprint category identification in step S4, including one of three categories: ball-hitting sound events, whistle events, and cheering events. The timestamp `t_evt` comes directly from the video and audio synchronization timeline, where B is the two-dimensional bounding box in the image coordinate system, and `p_evt` is the final confidence score. The event labeling results are written to the output buffer immediately after generation and can be read by the upper-layer interface for use in highlight replay or automatic event summary generation.

[0062] Please see Fig. 2 The left side is a simplified top view, marking the real sound source (referee) and the location of the mobile phone. The solid green line represents the direct sound path, the dashed green line represents the sound path reflected from the ground, and the solid yellow line represents the sound path reflected from the wall. There are gray dots representing mirrored sound sources on the reflection paths. The right side is the GCC-PHAT function graph, with red representing the real peak and blue representing the spurious reflection peak. The arrow at the top indicates the suppression weight, demonstrating the environmental adaptability of using visual priors to suppress reverberation interference.

[0063] Please see Fig. 3The system visually identifies the "potential event triggering area" (player); the spatial pointing data of the sound source is back-projected into the "sound source pointing line" in the image; by calculating the spatiotemporal overlap between the sound source pointing trajectory and the player's visual movement trajectory (bottom superimposed trajectory), the sound source-target association is realized, generating a "spatiotemporal consistency" metric for the core discrimination model to improve the event triggering confidence.

[0064] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is not limited by the foregoing description. Thus, all changes falling within the meaning and scope of the equivalents of the application are intended to be included within the scope of the invention.

[0065] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for automatic highlight event labeling based on sound source localization and visual fusion, characterized in that, The method applied to a mobile terminal including a microphone array and a camera comprises the following steps: Step S1: obtaining original acoustic data and original visual data, extracting an event acoustic feature sequence in the original acoustic data, and determining scene prior information based on the original visual data; Step S2: constructing an acoustic event feature group based on the original acoustic data and the event acoustic feature sequence, and determining sound source direction information according to the microphone array; Step S3: obtaining posture data of the mobile terminal, and constructing an array posture matrix; correcting the sound source direction information by using the array posture matrix to generate sound source space pointing data; Step S4: constructing a space fusion feature group by using the original visual data, the scene prior information and the sound source space pointing data, and identifying a spatial position of a sound event triggered in a target scene region to form an event trigger confidence; Step S5: determining a highlight event marking result according to the event trigger confidence.

2. The automatic highlight event tagging method based on sound source localization in visual fusion according to claim 1, wherein, Before obtaining the original acoustic data and the original visual data in step S1, the following steps are further included: starting the microphone array and the camera of the mobile terminal, and initializing collection of audio stream and video stream; loading a predefined acoustic event dictionary into the memory of the mobile terminal; initializing an inertial measurement unit of the mobile terminal, and constructing a spatial calibration relationship between the inertial measurement unit and the acoustic array.

3. The method for automatic highlight event labeling based on sound source localization and visual fusion according to claim 1, characterized in that, Step S1 comprises: collecting original acoustic data and original visual data by the microphone array and the camera of the mobile terminal respectively; performing event detection on the original acoustic data to identify an audio segment containing a potential sound event; performing feature matching and sparse coding on the audio segment based on the acoustic event dictionary, and extracting an event acoustic feature sequence; performing scene geometric structure analysis on the original visual data to identify main plane structures in the scene; performing three-dimensional space orientation and parameter fitting on the main plane structures to output a plane equation of the main plane structures in a geodetic coordinate system; extracting a boundary contour of the main plane structures, calculating a spatial scale, and constructing a three-dimensional geometric cage model describing a scene boundary; defining the plane equation and the three-dimensional geometric cage model as the scene prior information.

4. The method of claim 3, wherein, The feature matching and sparse coding on the audio segment based on the acoustic event dictionary comprises: calling multiple event sub-dictionaries related to sports events from the acoustic event dictionary, the event sub-dictionaries including a whistle sub-dictionary, a ball hitting sound sub-dictionary and a cheer sub-dictionary; converting the audio segment into a time-frequency feature representation to form an input signal to be processed; performing sparse decomposition on the input signal based on the acoustic event dictionary to extract sparse components representing different sound event types; identifying a sports event highlight type and a corresponding intensity contained in the audio segment according to a category attribute and an energy distribution of the sparse components; extracting an event component related to the sports event highlight from the sparse components; reconstructing an event acoustic feature signal sequence based on the event sub-dictionary and the event component.

5. The method for automatic highlight event labeling based on sound source localization and visual fusion according to claim 1, characterized in that, Step S2 comprises: calculating a time-domain energy envelope and a frequency-domain spectrum feature of a sound event based on the event acoustic feature sequence to form an initial event feature vector; performing multi-channel joint analysis on the original acoustic data to determine time difference information of sound waves arriving at different microphones; analyzing a candidate set of sound source directions based on the time difference information, the candidate set including multiple direction candidates formed by direct sound and reflected sound; fuse the initial event feature vector and the candidate set of sound source directions to form an acoustic event feature group; determine the sound source direction information according to the confidence weight of each candidate direction in the acoustic event feature group.

6. The method for automatic highlight event labeling based on sound source localization and visual fusion according to claim 1, characterized in that, The step S3 includes the following steps: acquire the attitude data of the mobile terminal, and construct an array attitude matrix, including: real-time acquire the three-axis acceleration data and three-axis angular velocity data of the mobile terminal body in the earth coordinate system through the inertial measurement unit of the mobile terminal; calculate the tilt angle of the mobile terminal relative to the gravity direction based on the three-axis acceleration data, and determine the pitch angle and roll angle; calculate the compensation dynamic offset according to the three-week angular velocity data, and correct the azimuth angle to output the yaw angle; calculate the rotation matrix around the Z-axis, Y-axis and X-axis based on the pitch angle, roll angle and yaw angle respectively, and generate a complete rotation transformation matrix; 7. The method for automatic highlight event labeling based on sound source localization and visual fusion according to claim 1, characterized in that, determine the complete rotation transformation matrix as the array attitude matrix for sound source space direction correction. The step S3 includes the following steps: acquire the static geometric layout coordinates of the microphone array in the mobile terminal body coordinate system; perform spatial transformation on the static geometric layout coordinates by using the array attitude matrix to generate dynamic spatial coordinates of the microphone array in the earth coordinate system; 8. The method of claim 7, wherein, perform coordinate system conversion on the sound source direction information based on the dynamic spatial coordinates to generate sound source space direction data independent of the mobile terminal attitude. The step S3 includes the following steps: perform matrix multiplication operation on the static geometric layout coordinates and the array attitude matrix to calculate the preliminary spatial coordinates of the microphone array in the earth coordinate system at the current time; 9. The method for automatic highlight event labeling based on sound source localization and visual fusion according to claim 1, wherein, perform translation compensation on the preliminary spatial coordinates based on the real-time displacement data of the mobile terminal to output the dynamic spatial coordinates of the microphone array. The step S4 includes the following steps: perform image recognition on the original visual data to detect and locate one or more potential event trigger regions in the target scene; project the sound source space direction data onto the image coordinate system formed by the original visual data to generate a sound source direction projection line; identify the specific sound event type corresponding to the event soundprint sequence; determine whether the spatial position pointed by the sound source direction projection line is consistent with the semantic type of the potential event trigger region based on the scene prior information; track the dynamic changes of the sound source space direction data within a continuous time period to generate a sound source direction trajectory; calculate the spatiotemporal consistency measure between the sound source direction trajectory and the visual motion trajectory of the moving target in the potential event trigger region; combine the spatial relationship measure, the visual features of the potential event trigger region, and the stability indicators of the sound source space direction data to construct a spatial fusion feature group; input the spatial fusion feature group into the pre-trained event positioning discriminant model to calculate the probability that each potential event trigger region is the real sound event occurrence position; 10. The method for automatic highlight event labeling based on sound source localization and visual fusion according to claim 1, characterized in that, determine the target scene region and the corresponding event trigger confidence based on the probability. The step S5 includes the following steps: compare the event trigger confidence with the preset confidence threshold; when the event trigger confidence is greater than or equal to the confidence threshold, determine that an effective highlight event occurs; acquire the target scene region and the timestamp information corresponding to the effective highlight event; Based on the target scene region and the timestamp information, a highlight event marking result is generated and output.

Citation Information

Patent Citations

  • Method based on audio / video combination for detecting highlight events in football video

    CN101650722A

  • Bird flock identification method and system based on ultra-high-definition video

    CN120012031A

  • Urban rail transit emergency processing method and system based on visual large model

    CN120047907A

  • Acoustic event positioning method, device and system and computer readable storage medium

    CN120370260A

  • Double-target automatic tracking shooting method and automatic tracking shooting device

    CN120614518A