Intelligent database video retrieval method based on big data technology

By performing spatiotemporal slicing and multimodal feature extraction on the video stream, combined with dynamic weight fusion and hierarchical hybrid index structure, the problem of difficulty in capturing deep semantic information and insufficient utilization of multimodal information in the existing video retrieval methods is solved, and efficient and accurate video retrieval and optimized system response speed is achieved.

CN120216722AActive Publication Date: 2025-06-27HEBEI ZHENGTONG ARCHIVES MANAGEMENT CO LTD

Patent Information

Application Number
CN202510287157.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Existing video retrieval methods are difficult to effectively capture the deep semantic information of video content, resulting in low correlation and accuracy of search results, and it is difficult to make full use of multimodal information in video data, affecting matching accuracy; at the same time, the system response speed is slow and it is difficult to meet the needs of large-scale video data processing.

Method used

By performing spatiotemporal slicing of the original video stream, visual, audio and text features are extracted, cross-modal correlation matrix is ​​constructed, and semantic feature vectors are generated using dynamic weight fusion method; at the same time, a hierarchical hybrid index structure is constructed based on the video heat value, so as to achieve rapid retrieval of high-hot videos and efficient storage of low-hot videos; using improved cosine similarity algorithm and user behavior data feedback, the accuracy and response speed of the search results are optimized.

Benefits of technology

Effectively integrate the multimodal features of video to improve matching accuracy and search accuracy; optimize the index structure, improve system response speed and search efficiency; make the search results more in line with user needs and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216722A_ABST
    Figure CN120216722A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent information retrieval, in particular to an intelligent database video retrieval method based on a big data technology, which comprises the following steps: S1, carrying out space-time slicing processing on an original video stream; s2, extracting a visual feature vector, an audio waveform vector and a text description vector; s3, constructing a cross-modal incidence matrix; s4, constructing a layered mixed index structure which comprises a real-time updating layer and a static storage layer; s5, after a user retrieval request is received, candidate video set screening is carried out based on the layered mixed index structure, and a retrieval result is generated; and S6, updating the cross-modal incidence matrix, and synchronously updating the weight parameter of the layer. According to the method, the retrieval accuracy is improved through cross-modal feature fusion, the data storage and retrieval efficiency is optimized by adopting a hierarchical mixed index structure, and the index weight is dynamically adjusted in combination with user behavior feedback, so that efficient, accurate and intelligent video retrieval is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent information retrieval, and particularly to an intelligent database video retrieval method based on big data technology. Background Art

[0002] With the rapid growth of video data, intelligent database video retrieval technology has become an important research direction in the field of big data processing; existing video retrieval methods mainly rely on keyword matching or index structures based on single-modal features, which are difficult to effectively capture the deep semantic information of video content, resulting in low relevance and accuracy of retrieval results; in addition, video data usually contains multi-modal information such as vision, audio, and text, and there are complex correlation relationships between different modalities, while traditional retrieval methods are difficult to make full use of these multi-modal information for effective matching; in addition, due to the huge amount of video data, how to improve the system response speed while ensuring retrieval accuracy has also become an urgent problem to be solved in the field of intelligent database video retrieval.

[0003] The following difficulties generally exist in the prior art during video retrieval: First, the fusion method of multi-modal information lacks an effective dynamic weight adjustment mechanism, resulting in unbalanced contributions of different modalities to the retrieval results and affecting the matching accuracy; second, the design of the index structure is difficult to balance the rapid update of high-popularity videos and the long-term storage of low-popularity videos, affecting the retrieval efficiency; third, the operation behavior data of users cannot be effectively fed back into the retrieval system, resulting in the retrieval results unable to be optimized according to user preferences and affecting the user experience. Therefore, there is an urgent need for an intelligent database video retrieval method based on big data technology to solve the above problems. Summary of the Invention

[0004] Based on the above purpose, the present invention provides an intelligent database video retrieval method based on big data technology.

[0005] An intelligent database video retrieval method based on big data technology includes the following steps:

[0006] S1: Perform spatio-temporal slicing processing on the original video stream to generate video segment units with timestamp marks;

[0007] S2: Extract visual feature vectors, audio waveform vectors, and text description vectors from the video segment units generated in S1 to generate a multi-modal feature vector group;

[0008] S3: Construct a cross-modal association matrix based on user historical behavior data, and perform dynamic weight fusion on the multi-modal feature vector group to generate a semantic feature vector;

[0009] S4: construct a hierarchical hybrid index structure according to the video heat value, wherein the hierarchical hybrid index structure includes a real-time update layer and a static storage layer, and then write the semantic feature vector generated by S3 into the real-time update layer or the static storage layer;

[0010] S5: After receiving the user's search request, the candidate video set is screened based on the hierarchical hybrid index structure, and the matching degree between the semantic feature vector of the candidate video set and the search request is calculated by the improved cosine similarity algorithm to generate the search results, and the user's operation behavior data on the search results is recorded at the same time;

[0011] S6: Update the cross-modal association matrix based on the recorded user operation behavior data, and synchronously trigger the weight parameters of the real-time update layer in the hierarchical hybrid index structure to optimize the accuracy of subsequent retrieval results.

[0012] Optionally, the S1 specifically includes:

[0013] S11: Detecting scene switching points based on the color histogram difference between video frames, calculating the three-dimensional histogram of the HSV color space for consecutive video frames, and determining the scene switching point when the difference in Bhattacharyya coefficients between adjacent frames exceeds a set threshold;

[0014] S12: calculating the sum of the gradient amplitudes of each frame within a range of N frames before and after the detected scene switching point, and selecting the frame with the largest gradient sum as the key frame, where N is a preset sliding window size;

[0015] S13: With the key frame as the center, the time window is expanded to both sides. The initial time window length is set to T seconds. The time window boundary is dynamically adjusted according to the scene switching point spacing on both sides of the key frame. The specific adjustments include:

[0016] If the distance between the left adjacent scene switching point and the current key frame is less than T / 2, the left boundary is adjusted to the position of the left scene switching point;

[0017] If the distance between the right adjacent scene switching point and the current key frame is less than T / 2, the right boundary is adjusted to the position of the right scene switching point;

[0018] S14: intercepting the video stream in the adjusted time window, generating a video segment unit with a start timestamp and an end timestamp, controlling the timestamp accuracy at the millisecond level, and marking the frame number index of the key frame to which it belongs.

[0019] Optionally, the S2 specifically includes:

[0020] S21: The video frame data in each video clip unit generated by S1 is processed using a ResNet-50 network structure, and after normalizing, resizing, and data enhancement of the video frame, a visual feature vector V of a fixed dimension is extracted;

[0021] S22: For the audio data within the video clip unit, first perform frame division on the audio signal using the short-time Fourier transform, and then calculate the 13-dimensional coefficients of each frame of audio using the Mel-frequency cepstral coefficient extraction algorithm, thereby obtaining the audio waveform vector M;

[0022] S23: For the text information within the video clip unit, first use optical character recognition technology to recognize and extract the text information in the image, and then perform semantic encoding on the recognized text to obtain a text description vector W with a fixed dimension;

[0023] S24: First, perform normalization processing on the visual feature vector, audio waveform vector, and text description vector obtained from S21, S22, and S23 respectively; then perform vector splicing according to a predetermined dimension ratio to generate a multi-modal feature vector group.

[0024] Optionally, S24 specifically includes:

[0025] S241: Perform L2 normalization processing on the visual feature vector, audio waveform vector, and text description vector obtained from S21, S22, and S23 respectively to obtain the normalized visual feature vector V′, audio waveform vector M′, and text description vector W′;

[0026] S242: Assume that the total dimension of the finally generated multi-modal feature vector group is D, and determine the dimension allocation of each modal feature vector according to a ratio of 5:3:2;

[0027] S243: Perform dimension adjustment on the normalized visual feature vector V′, audio waveform vector M′, and text description vector W′ to make them match the target dimension;

[0028] S244: Splice the adjusted visual feature vector V″, audio waveform vector M″, and text description vector W″ in a sequential manner to generate the final multi-modal feature vector group F.

[0029] Optionally, S3 specifically includes:

[0030] S31: Perform statistical processing on the user's historical behavior data. Assume that in the k-th record of the user's historical behavior data, the user interaction scores collected for the visual, audio, and text modalities are denoted as x V,k , x M,k and x W,k , and calculate the average scores of each modality, which are respectively denoted as and Furthermore, based on the user interaction scores between modalities, construct a cross-modal correlation matrix A using the Pearson correlation coefficient;

[0031] S32: Calculate the weight parameters of each modality according to the correlation coefficients among vision, audio, and text in the cross-modal correlation matrix A. Among them, the weight parameter α of the visual modality is calculated as follows: The weight parameter β of the audio modality is calculated as follows: The weight parameter γ of the text modality is calculated as follows: The sum of all weight parameters is 1.

[0032] S33: Use the dynamic weight coefficients to perform weighted fusion on the feature vectors of each modality to generate the semantic feature vector S. The formula is: S = αV″ + βM″ + γW″, where S is the finally generated semantic feature vector.

[0033] Optionally, the specific content of S4 includes:

[0034] S41: For each video clip unit, calculate the video popularity value H according to the fixed weight.

[0035] S42: Set a preset threshold T, and compare the calculated H with T to determine the belonging index layer. The calculation formula for the index layer identifier I is: Among them, I = 1 means that the video clip unit is classified into the real-time update layer, and I = 0 means it is classified into the static storage layer;

[0036] S43: Construct a hierarchical hybrid index structure including a real-time update layer and a static storage layer. The real-time update layer supports incremental write operations, and the static storage layer is used to store video clip units belonging to low popularity.

[0037] S44: According to the value of I, write the semantic feature vector into the corresponding index layer.

[0038] Optionally, the specific content of S5 includes:

[0039] S51: Receive the retrieval request input by the user and convert the retrieval request into a text description vector.

[0040] S52: Based on the constructed hierarchical hybrid index structure, perform retrievals on the real-time update layer and the static storage layer respectively. According to the video popularity value, timestamp, and preset index screening conditions, select the candidate video sets that meet the conditions from each layer.

[0041] S53: Extract the semantic feature vectors stored in each video clip unit in the candidate video set, calculate the matching degree with the semantic query vector corresponding to the retrieval request using an improved cosine similarity algorithm, then sort according to the similarity value, and screen out the video clip with the highest matching degree to generate the final retrieval result.

[0042] S54: Record the user's click, play, like, and comment behaviors and store them in the historical behavior dataset.

[0043] Optionally, the S52 specifically includes:

[0044] S521: Define a set of conditions for screening candidate video segment units, including the video popularity value threshold T H , the timestamp range [T min , T max , and the preset index screening condition C;

[0045] S522: For each video segment unit stored in the index structure, calculate its screening score R;

[0046] S523: According to the screening score R calculated in S522, screen out the video segment units that meet the conditions in the real-time update layer and the static storage layer. If S≥T S , then this video segment unit is added to the candidate video set, where T S is the preset screening threshold.

[0047] Optionally, the S53 specifically includes:

[0048] S531: Extract the semantic feature vectors stored in each video segment unit from the candidate video set and convert the user input retrieval request into a semantic query vector;

[0049] S532: Use an improved cosine similarity algorithm to calculate the matching degree between the semantic feature vector of each candidate video segment unit and the semantic query vector. This improved cosine similarity algorithm measures the semantic relevance between each candidate video segment and the retrieval request by introducing dynamic weight adjustment to eliminate the influence differences between different modalities;

[0050] S533: Sort all candidate video segment units in descending order of matching degree;

[0051] S534: Select the video segment unit with the highest matching degree in the sorting result according to the predetermined return quantity as the final retrieval result and return it to the user.

[0052] Optionally, the S6 specifically includes:

[0053] S61: Conduct statistical analysis on the recorded user click, play, like, and comment behavior data, classify the interaction scores for visual, audio, and text modalities within each video segment unit respectively, and calculate the latest correlation coefficient between modalities according to the weighted statistical method;

[0054] S62: Update the cross-modal association matrix using the latest correlation coefficient, replacing the parameters reflecting the correlation coefficients between the visual, audio, and text modalities in the original matrix with the determined values obtained through statistical analysis and calculation to form a new cross-modal association matrix;

[0055] S63: According to the correlation coefficients of each modality in the updated cross-modal association matrix, update the weight parameters corresponding to each modality in the layer in real time according to the predetermined mapping relationship, and synchronously transmit the corresponding weight parameters to the real-time update layer in the hierarchical hybrid index structure to trigger the immediate adjustment of the weight parameters in the real-time update layer, ensuring that the weight parameters are consistent with the latest user behavior data during subsequent writing and retrieval processes.

[0056] Advantages of the present invention:

[0057] In the present invention, by constructing a cross-modal association matrix and adopting a dynamic weight fusion method, the visual, audio, and text features of the video can be effectively integrated, improving the matching accuracy of multi-modal information; through the improved cosine similarity algorithm and combined with the dynamic weight adjustment mechanism, the contribution ratio of each modality feature during the retrieval process is more reasonable, thereby improving the precision of video retrieval; at the same time, the present invention updates the cross-modal association matrix based on the user's historical behavior data, enabling the system to continuously optimize the feature fusion method and making the retrieval results more in line with the user's needs.

[0058] In the present invention, by constructing a hierarchical hybrid index structure, the video data is classified and stored according to the popularity value, realizing the rapid retrieval of high-popularity videos and the efficient storage of low-popularity videos, and adopting an incremental update mechanism in the real-time update layer to improve the retrieval response speed; in addition, through the feedback of user behavior data, the weight parameters in the index structure are dynamically adjusted, enabling the system to optimize the subsequent retrieval strategy according to the user's actual interaction behavior, thereby improving the intelligent level of video retrieval. Description of the Drawings

[0059] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0060] Figure 1 Schematic diagram of the intelligent database video retrieval method according to the embodiment of the present invention;

[0061] Figure 2 Schematic diagram of the process of generating retrieval results according to the embodiment of the present invention. Detailed Embodiments

[0062] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specifically describing the embodiments, and are not intended to specifically limit the present invention.

[0063] It should be noted that in the specification, references to "one embodiment", "an embodiment", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when combining embodiments to describe a specific feature, structure, or characteristic, implementing such feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0064] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that may not be explicitly described.

[0065] As Figure 1 - Figure 2 shown, an intelligent database video retrieval method based on big data technology includes the following steps:

[0066] S1: Perform spatio-temporal slicing on the original video stream to generate video segment units with timestamp markings;

[0067] S2: Extract visual feature vectors, audio waveform vectors, and text description vectors from the video segment units generated in S1 to generate a multi-modal feature vector group;

[0068] S3: Construct a cross-modal association matrix based on user historical behavior data, and perform dynamic weight fusion on the multi-modal feature vector group to generate semantic feature vectors;

[0069] S4: Construct a hierarchical hybrid index structure according to the video popularity value. The hierarchical hybrid index structure includes a real-time update layer and a static storage layer, and then write the semantic feature vectors generated in S3 into the real-time update layer or the static storage layer, where the real-time update layer supports incremental write operations;

[0070] S5: After receiving the user's retrieval request, filter the candidate video set based on the hierarchical hybrid index structure, and calculate the matching degree between the semantic feature vector of the candidate video set and the retrieval request through an improved cosine similarity algorithm to generate the retrieval result. At the same time, record the operation behavior data of the user on the retrieval result;

[0071] S6: Update the cross-modal association matrix according to the recorded user operation behavior data, and synchronously trigger the weight parameters of the real-time update layer in the hierarchical hybrid index structure to optimize the accuracy of subsequent retrieval results.

[0072] S1 specifically includes:

[0073] S11: Detect the scene switching point based on the difference degree of the color histograms between video frames. Calculate the three-dimensional histogram of the HSV color space for consecutive video frames. When the Bhattacharyya coefficient difference between adjacent frames exceeds the set threshold, it is determined as the scene switching point;

[0074] The specific calculation formula for the Bhattacharyya coefficient difference is: where D Bhattacharyya is the Bhattacharyya coefficient difference degree; H1(i) and H2(i) respectively represent the probability densities of adjacent frames in the i-th dimensional histogram distribution; n is the histogram dimension. In this solution, the HSV color space includes 16 intervals each for hue H, saturation S, and value V, that is, n = 16×16×16; when D Bhattacharrya > 0.3, it is determined that this frame is the scene switching point;

[0075] S12: Calculate the total gradient magnitude of each frame within N frames before and after the detected scene switching point, and select the frame with the largest total gradient as the key frame, where N is the preset sliding window size; the calculation formula for the total gradient magnitude is as follows: where G sum is the total gradient magnitude of the current frame; I is the grayscale value of the frame image; x, y are pixel coordinates; and are the horizontal and vertical gradients calculated by the Sobel operator respectively; the above-mentioned sliding window size N is set according to the video frame rate FPS, and the formula is: where represents rounding down;

[0076] S13: With the key frame as the center, expand the time window to both sides. The initial time window length is set to T seconds, and the time window boundary is dynamically adjusted according to the scene switching point spacing on both sides of the key frame. The specific adjustment includes:

[0077] If the distance between the left adjacent scene switching point and the current key frame is less than T / 2, adjust the left boundary to the position of the left scene switching point;

[0078] If the distance between the right adjacent scene transition point and the current key frame is less than T / 2, adjust the right boundary to the position of the right scene transition point;

[0079] S14: Intercept the video stream within the adjusted time window to generate a video segment unit with a start timestamp and an end timestamp, control the timestamp accuracy at the millisecond level, and mark the frame number index of the associated key frame; where the timestamp is in the ISO 8601 extended format and is recorded in the form of YYYY-MM-DDThh:mm:ss.sssZ; the index is the globally unique identifier of the key frame associated with each segment unit, and the format is <video ID>_<key frame sequence number>.

[0080] S2 specifically includes:

[0081] S21: Process the video frame data within each video segment unit generated in S1 using the ResNet-50 network structure, and after normalizing, standardizing the size, and data augmentation of the video frames, extract a visual feature vector V with a fixed dimension;

[0082] The specific steps are as follows:

[0083] Perform normalization processing on the video frames to map the pixel value range to 0, 1. The normalization formula is as follows: where P′ represents the pixel value of the normalized video frame, P is the original pixel value, P min and P max represent the minimum and maximum pixel values respectively;

[0084] Perform size standardization on the normalized frames to make all input frame sizes consistent. Assume the input frame size is W×H, and after standardization, it is uniformly adjusted to W′×H′. The adjustment formula is as follows: W′ = W×S and H′ = H×S, where S represents the scaling ratio coefficient, and its calculation formula is: where W t and H t are the standard size width and height respectively;

[0085] Use the ResNet-50 network for feature extraction. Input the standardized frames into the convolutional layer of ResNet-50 to output a visual feature vector V with a fixed dimension. The expression is: V = F ResNet (I), where I is the input frame image matrix, and F ResNet (·) represents the ResNet-50 network processing function, and V is the finally extracted visual feature vector.

[0086] S22: For the audio data within the video clip unit, first perform frame division on the audio signal using the Short-Time Fourier Transform (STFT), and then calculate the 13-dimensional coefficients of each audio frame using the Mel Frequency Cepstral Coefficient (MFCC) extraction algorithm, thereby obtaining the audio waveform vector M;

[0087] The specific steps include:

[0088] First, perform STFT transformation on the input audio signal to obtain the spectral data S(f, t), and its calculation formula is as follows: where S(f, t) is the time-frequency matrix, x(n) represents the discrete audio signal, w(n) is the window function, N is the window length, f is the frequency index, and t is the time frame index;

[0089] Then, calculate the Mel Frequency Cepstral Coefficient (MFCC), and the formula is as follows:

[0090] where M k is the k-th dimensional MFCC feature, H k (f) is the weight function of the Mel filter bank, and F is the number of filters.

[0091] S23: For the text information within the video clip unit, first use the Optical Character Recognition (OCR) technology to recognize and extract the text information in the image, and then perform semantic encoding on the recognized text to obtain a text description vector W with a fixed dimension;

[0092] S24: First, perform normalization processing on the visual feature vector, audio waveform vector, and text description vector obtained from S21, S22, and S23 respectively; then perform vector splicing according to a predetermined dimension ratio to generate a multi-modal feature vector group, and append the time stamp index of the source video clip unit; Through the above steps and solutions, the feature extraction, normalization, and fixed-ratio splicing of each modal data achieve the precise quantitative expression of multi-modal features, ensuring the efficient fusion of data at the same scale, thereby improving the accuracy and response speed of the video retrieval system.

[0093] S24 specifically includes:

[0094] S241: Perform L2 normalization processing on the visual feature vector, audio waveform vector, and text description vector obtained from S21, S22, and S23 respectively to obtain the normalized visual feature vector V′, audio waveform vector M′, and text description vector W′. The normalization calculation formula is as follows: where ||·|| represents the calculation of the Euclidean norm;

[0095] S242: Let the total dimension of the finally generated multi-modal feature vector group be D. Determine the dimension allocation of each modal feature vector according to the ratio of 5:3:2. The calculation formula is as follows: and D W = D - (D V + D M ), where D V , D M and D W represent the target dimensions of visual features, audio waveform features, and text description features respectively. represents the floor operation to ensure that the sum of the vector dimensions is consistent with D;

[0096] S243: Perform dimension adjustment on the normalized visual feature vector V′, audio waveform vector M′, and text description vector W′ to make them match the target dimensions. Specifically, use a linear projection transformation matrix for dimensionality reduction. The calculation formulas are as follows: V″ = W V V′, M″ = W M M′, and W″ = W W W′, where and are the dimensionality reduction projection matrices for visual, audio, and text features respectively. V″, M″, and W″ are the target dimension feature vectors after adjustment;

[0097] S244: Concatenate the adjusted visual feature vector V″, audio waveform vector M″, and text description vector W″ in a sequential manner to generate the final multi-modal feature vector group F. The expression is: F = [V″; M″; W″], where [·; ·; ·] represents the concatenation operation of vectors, making the dimension of the final feature vector group F be D; Through the above steps, the unified normalization processing, ratio allocation, dimension adjustment, and concatenation of visual, audio, and text feature vectors are realized, ensuring the structural consistency of the multi-modal feature vector group, and optimizing the subsequent video retrieval process by combining the timestamp index.

[0098] S3 specifically includes:

[0099] S31: Perform statistical processing on the user historical behavior data. In the k-th record of the user historical behavior data, the user interaction scores collected for visual, audio, and text modalities are denoted as x V,k , x M,k and x W,k (where k = 1, 2,..., L, and L is the total number of records), and calculate the average scores of each modality, denoted as and respectively. Then, based on the user interaction scores between modalities, construct a cross-modal correlation matrix A using the Pearson correlation coefficient. The correlation coefficient r VM between vision and audio is calculated as follows:

[0100] The correlation coefficient r between vision and text WW and the correlation coefficient r between audio and text MW are calculated in the same form respectively; the expression of the cross-modal correlation matrix A is: wherein, the symmetric relationship holds, that is, r VM = e MV , r VW = r WV and r MW = r WM ; each element in the matrix is a determined numerical value;

[0101] S32: According to the correlation coefficients between vision, audio and text in the cross-modal correlation matrix A, calculate the weight parameters of each modality; among them, the weight parameter α of the vision modality is calculated as follows: The weight parameter β of the audio modality is calculated as follows: The weight parameter γ of the text modality is calculated as follows: The sum of each weight parameter is 1;

[0102] S33: Based on the vision feature vector V″, audio waveform vector M″ and text description vector W″ obtained in S24 after L2 normalization and dimension adjustment, use the dynamic weight coefficient to perform weighted fusion on each modality feature vector to generate the semantic feature vector S. The formula is: S = αV″ + βM″ + γW″, where S is the finally generated semantic feature vector, and each of its components is a determined numerical value; through the above steps, the constructed cross-modal correlation matrix can accurately reflect the internal correlation between each modality in the user's historical behavior data, and the dynamic weight fusion mechanism ensures the organic integration of multi-modal features, thereby improving the accuracy and response speed of video retrieval.

[0103] S4 specifically includes:

[0104] S41: For each video segment unit, calculate the video popularity value H according to the fixed weight; its calculation formula is: H = aV1 + bL + cC, where V1 represents the number of views of the video, L represents the number of likes of the video, and C represents the number of comments of the video; a, b and c are positive definite coefficients determined by statistics and are all fixed numerical values;

[0105] S42: Set a preset threshold T, and compare the calculated H with T to determine the belonging index layer; define the calculation formula of the index layer identifier I as: wherein, I = 1 means the video segment unit is classified into the real-time update layer, and I = 0 means it is classified into the static storage layer;

[0106] S43: Construct a hierarchical hybrid index structure including a real-time update layer and a static storage layer; the real-time update layer supports incremental write operations, and the static storage layer is used to store video segment units belonging to low popularity.

[0107] S44: According to the value of I, write the semantic feature vector into the corresponding index layer; through the above steps, a hierarchical hybrid index structure is constructed based on the video popularity value, and the semantic feature vectors generated in S3 are classified and written into the real-time update layer or the static storage layer according to the popularity value, thereby improving the real-time response ability and data management efficiency of the video retrieval system.

[0108] S5 specifically includes:

[0109] S51: Receive the retrieval request input by the user and convert the retrieval request into a text description vector.

[0110] S52: Based on the constructed hierarchical hybrid index structure, retrieve the real-time update layer and the static storage layer respectively, and select candidate video sets that meet the conditions from each layer according to the video popularity value, timestamp, and preset index screening conditions.

[0111] S53: Extract the semantic feature vectors stored in each video segment unit in the candidate video set, calculate the matching degree with the semantic query vector corresponding to the retrieval request using an improved cosine similarity algorithm, then sort according to the similarity value, and screen out the video segment with the highest matching degree to generate the final retrieval result; the improved cosine similarity algorithm explicitly introduces a dynamic weight adjustment parameter on the basis of the traditional cosine similarity to eliminate the influence differences between various modal features and achieve accurate comparison of various modal semantic information.

[0112] S54: Record the user's click, play, like, and comment behaviors and store them in the historical behavior data set to provide a basis for the cross-modal correlation matrix update in the subsequent step S6; the above steps of text vectorization of the user's retrieval request, screening of candidate video sets based on the hierarchical hybrid index structure, and matching and sorting of the improved cosine similarity algorithm effectively improve the accuracy and response speed of video retrieval, ensuring the relevance and timeliness of the retrieval results.

[0113] S52 specifically includes:

[0114] S521: Define a set of conditions for screening candidate video segment units, including the video popularity value threshold T H , the timestamp range [T min , T max , and the preset index screening condition C; among them, the video popularity value threshold T H is used to screen high-popularity videos; the timestamp range [T min , T maxLimit the video release or update time to be within this range; the preset index filtering condition C includes video category, user preference tags, and other specific retrieval restrictions;

[0115] S522: For each video segment unit stored in the index structure, calculate its screening score R, and the calculation formula is as follows: R = w H ·H + w T ·R T + w C ·R C , where H is the popularity value of the video segment, and its calculation method is based on step S41; R T is the timestamp correlation score, and the calculation formula is as follows:

[0116] where T video represents the timestamp of the video segment, ensuring that its time correlation score is within the range of 0, 1; R C is the index filtering condition matching degree, and the calculation formula is as follows: where C video is the label set of the video segment, |C| represents the number of labels in the filtering condition, and |C ∩ C video | represents the number of labels that meet the filtering condition; w H , w T and w C are the weight coefficients of video popularity, timestamp correlation, and index filtering conditions respectively, and their sum satisfies: w H + w T + w C = 1;

[0117] S523: According to the screening score R calculated in S522, screen out the video segment units that meet the conditions in the real-time update layer and the static storage layer. If S ≥ T S , then this video segment unit is added to the candidate video set, where T S is the preset screening threshold to ensure that only the most relevant video segments are selected as the candidate video set; the above steps achieve the precise screening of the candidate video set through the quantitative calculation of video popularity values, timestamps, and preset index filtering conditions, effectively improving the accuracy of the retrieval candidate range and the response speed of the system.

[0118] S53 specifically includes:

[0119] S531: Extract the semantic feature vectors stored in each video segment unit from the candidate video set, and convert the retrieval request input by the user into a semantic query vector;

[0120] S532: Use an improved cosine similarity algorithm to calculate the matching degree between the semantic feature vector of each candidate video segment unit and the semantic query vector. This improved cosine similarity algorithm eliminates the influence differences between different modalities by introducing dynamic weight adjustment, thereby measuring the semantic relevance between each candidate video segment and the retrieval request;

[0121] S533: Sort all candidate video segment units in descending order of the matching degree, so that the video segments with higher matching degrees are ranked at the front;

[0122] S534: Select the video segment unit with the highest matching degree in the sorting result according to the predetermined return quantity as the final retrieval result and return it to the user.

[0123] The calculation process for generating the final retrieval result is as follows:

[0124] S531: For each video segment unit in the candidate video set, extract the stored semantic feature vector and denote it as S i ; At the same time, convert the user's retrieval request into a semantic query vector with a fixed dimension through a predetermined semantic encoding module and denote it as Q, where i is the index of the candidate video segment unit;

[0125] S532: Use an improved cosine similarity algorithm to calculate the matching degree between each candidate video segment unit and the semantic query vector. The formula is: where C i represents the improved cosine similarity between candidate video segment unit i and the semantic query vector; S ij and Q j are the j-th components of vectors S i and Q respectively; λ j is the weight coefficient corresponding to the j-th dimension, which is determined by the multi-modal feature dynamic weight fusion process and is a fixed constant; D is the total dimension of the semantic feature vector;

[0126] S533: Sort the similarities C i of all candidate video segment units in descending order, and according to the sorting result, select the top K candidate video segment units with the highest similarity as the final retrieval result according to the predetermined return quantity K and return it to the user.

[0127] S6 specifically includes:

[0128] S61: Conduct statistical analysis on the recorded user click, play, like, and comment behavior data, classify the interaction scores for the visual, audio, and text modalities within each video segment unit respectively, and calculate the latest correlation coefficient between each modality according to the weighted statistical method;

[0129] S62: Update the cross-modal association matrix using the latest correlation coefficient, and replace each parameter reflecting the correlation coefficients among the visual, audio, and text modalities in the original matrix with the determined values obtained through statistical analysis and calculation to form a new cross-modal association matrix;

[0130] S63: According to the correlation coefficients of each modality in the updated cross-modal association matrix, update the weight parameters corresponding to each modality in the layer in real time according to the predetermined mapping relationship, and synchronously transmit the corresponding weight parameters to the real-time update layer in the hierarchical hybrid index structure to trigger the immediate adjustment of the weight parameters in the real-time update layer, ensuring that the weight parameters are consistent with the latest user behavior data during subsequent writing and retrieval processes; The above steps achieve the real-time update of the cross-modal association matrix through the precise statistics and weighted calculation of user operation behavior data, and synchronously adjust the weight parameters in the real-time update layer, thereby ensuring the accurate response of the retrieval results to user needs and the continuous optimization of system performance.

[0131] The present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0132] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An intelligent database video retrieval method based on big data technology, characterized in that: The following steps are involved: S1: Perform spatiotemporal slicing on the original video stream to generate video clip units with timestamps; S2: extracting visual feature vectors, audio waveform vectors and text description vectors from the video clip units generated by S1 to generate a multimodal feature vector group; S3: Construct a cross-modal association matrix based on user historical behavior data, and dynamically weight the multi-modal feature vector group to generate a semantic feature vector; S4: construct a hierarchical hybrid index structure according to the video heat value, wherein the hierarchical hybrid index structure includes a real-time update layer and a static storage layer, and then write the semantic feature vector generated by S3 into the real-time update layer or the static storage layer; S5: After receiving the user's search request, the candidate video set is screened based on the hierarchical hybrid index structure, and the matching degree between the semantic feature vector of the candidate video set and the search request is calculated by the improved cosine similarity algorithm to generate the search results, and the user's operation behavior data on the search results is recorded at the same time; S6: Update the cross-modal association matrix based on the recorded user operation behavior data, and synchronously trigger the weight parameters of the real-time update layer in the hierarchical hybrid index structure to optimize the accuracy of subsequent retrieval results.

2. According to the intelligent database video retrieval method based on big data technology according to claim 1, it is characterized in that: The S1 specifically includes: S11: Detecting scene switching points based on the color histogram difference between video frames, calculating the three-dimensional histogram of the HSV color space for consecutive video frames, and determining the scene switching point when the difference in Bhattacharyya coefficients between adjacent frames exceeds a set threshold; S12: calculating the sum of the gradient amplitudes of each frame within a range of N frames before and after the detected scene switching point, and selecting the frame with the largest gradient sum as the key frame, where N is a preset sliding window size; S13: With the key frame as the center, the time window is expanded to both sides. The initial time window length is set to T seconds. The time window boundary is dynamically adjusted according to the scene switching point spacing on both sides of the key frame. The specific adjustments include: If the distance between the left adjacent scene switching point and the current key frame is less than T / 2, the left boundary is adjusted to the position of the left scene switching point; If the distance between the right adjacent scene switching point and the current key frame is less than T / 2, the right boundary is adjusted to the position of the right scene switching point; S14: intercepting the video stream in the adjusted time window, generating a video segment unit with a start timestamp and an end timestamp, controlling the timestamp accuracy at the millisecond level, and marking the frame number index of the key frame to which it belongs.

3. The intelligent database video retrieval method based on big data technology according to claim 1 is characterized in that: The S2 specifically includes: S21: The video frame data in each video clip unit generated by S1 is processed using a ResNet-50 network structure, and after normalizing, resizing, and data enhancement of the video frame, a visual feature vector V of a fixed dimension is extracted; S22: for the audio data in the video clip unit, firstly perform frame processing on the audio signal by using short-time Fourier transform, and then calculate the 13-dimensional coefficients of each frame of audio by using Mel-frequency cepstral coefficient extraction algorithm, so as to obtain the audio waveform vector M; S23: For the text information in the video clip unit, firstly use the optical character recognition technology to recognize and extract the text information in the image, and then perform semantic encoding on the recognized text to obtain a text description vector W of a fixed dimension; S24: first normalize the visual feature vectors, audio waveform vectors and text description vectors obtained in S21, S22 and S23 respectively; then concatenate the vectors according to a predetermined dimensional ratio to generate a multimodal feature vector group.

4. The intelligent database video retrieval method based on big data technology according to claim 3 is characterized in that: The S24 specifically includes: S241: performing L2 normalization processing on the visual feature vector, audio waveform vector and text description vector obtained in S21, S22 and S23 respectively, to obtain a normalized visual feature vector V′, an audio waveform vector M′ and a text description vector W′; S242: Assuming that the total dimension of the multimodal feature vector group finally generated is D, the dimension allocation of each modal feature vector is determined according to the ratio of 5:3:2; S243: adjusting the dimensions of the normalized visual feature vector V′, the audio waveform vector M′, and the text description vector W′ to match the target dimension; S244: The adjusted visual feature vector V″, audio waveform vector M″ and text description vector W″ are concatenated in a sequence to generate a final multimodal feature vector group F.

5. The intelligent database video retrieval method based on big data technology according to claim 4 is characterized in that: The S3 specifically includes: S31: Perform statistical processing on the user's historical behavior data. Suppose the user interaction scores collected for the visual, audio and text modalities in the kth record of the user's historical behavior data are recorded as x V,k ,x M,k and x W,k , and calculate the average score of each mode and record it as and Then, based on the user interaction scores between each modality, the Pearson correlation coefficient is used to construct the cross-modal correlation matrix A; S32: Calculate the weight parameters of each modality according to the correlation coefficients among vision, audio and text in the cross-modal association matrix A; wherein the visual modality weight parameter α is calculated as follows: The audio modal weight parameter β is calculated as follows: The text modality weight parameter γ is calculated as follows: The sum of all weight parameters is 1; S33: Using the dynamic weight coefficient, perform weighted fusion on the feature vectors of each modality to generate a semantic feature vector S. The formula is: S = αV″+βM″+γW″, where S is the final generated semantic feature vector.

6. The intelligent database video retrieval method based on big data technology according to claim 1 is characterized in that: The S4 specifically includes: S41: For each video clip unit, calculate the video heat value H according to a fixed weight; S42: Set a preset threshold T, and compare the calculated H with T to determine the index layer to which it belongs; define the calculation formula of the index layer identifier I as follows: Wherein, I=1 indicates that the video clip unit is classified into the real-time update layer, and I=0 indicates that it is classified into the static storage layer; S43: constructing a hierarchical hybrid index structure including a real-time update layer and a static storage layer; the real-time update layer supports incremental write operations, and the static storage layer is used to store video clip units with low heat; S44: According to the value of I, the semantic feature vector is written into the corresponding index layer.

7. The intelligent database video retrieval method based on big data technology according to claim 1 is characterized in that: The S5 specifically includes: S51: receiving a search request input by a user, and converting the search request into a text description vector; S52: Based on the constructed hierarchical hybrid index structure, the real-time update layer and the static storage layer are searched respectively, and according to the video heat value, timestamp and preset index screening conditions, a candidate video set that meets the conditions is selected from each layer; S53: extracting the semantic feature vector stored in each video clip unit in the candidate video set, and using an improved cosine similarity algorithm to calculate the matching degree with the semantic query vector corresponding to the search request, and then sorting according to the similarity value, screening out the video clip with the highest matching degree, and generating the final search result; S54: Record the user's click, play, like and comment behaviors, and store them in a historical behavior data set.

8. The intelligent database video retrieval method based on big data technology according to claim 7 is characterized in that: The S52 specifically includes: S521: Define a set of conditions for screening candidate video clip units, including a video heat value threshold T H , timestamp range [T min , T max ] and preset index screening condition C; S522: For each video clip unit stored in the index structure, calculate its screening score R; S523: According to the screening score R calculated in S522, screen out the video clip units that meet the conditions in the real-time update layer and the static storage layer. If S≥T S , then the video segment unit is added to the candidate video set, where T S The preset filtering threshold.

9. The intelligent database video retrieval method based on big data technology according to claim 8 is characterized in that: The S53 specifically includes: S531: extracting the semantic feature vector stored in each video clip unit from the candidate video set, and converting the search request input by the user into a semantic query vector; S532: using an improved cosine similarity algorithm to calculate the matching degree between the semantic feature vector of each candidate video segment unit and the semantic query vector. The improved cosine similarity algorithm eliminates the influence difference between different modalities by introducing dynamic weight adjustment, thereby measuring the semantic relevance between each candidate video segment and the retrieval request. S533: sorting all candidate video segment units from high to low according to the matching degree; S534: Select the video clip unit with the highest matching degree in the sorted results according to the predetermined return quantity as the final search result, and return it to the user.

10. The intelligent database video retrieval method based on big data technology according to claim 1, characterized in that: The S6 specifically includes: S61: Statistically analyzing the recorded user click, play, like and comment behavior data, classifying the interaction scores for visual, audio and text modalities in each video clip unit, and calculating the latest correlation coefficient between each modality based on a weighted statistical method; S62: using the latest correlation coefficient to update the cross-modal association matrix, replacing various parameters in the original matrix reflecting the correlation coefficients between the visual, audio and textual modalities with determined values ​​calculated through statistical analysis, so as to form a new cross-modal association matrix; S63: Based on the correlation coefficients of each modality in the updated cross-modal association matrix, the weight parameters corresponding to each modality in the layer are updated in real time according to the predetermined mapping relationship, and the corresponding weight parameters are synchronously transmitted to the real-time update layer in the hierarchical hybrid index structure, triggering instant adjustment of the weight parameters in the real-time update layer to ensure that the weight parameters are consistent with the latest user behavior data during subsequent writing and retrieval.

Citation Information

Patent Citations

  • Multimodal recommendation method based on hypergraph collaborative filtering

    CN117909599A

  • Mechanical equipment fault diagnosis method and system adopting correlation analysis

    CN118939988A

  • Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device

    CN119027775A

  • Multi-modal standardized knowledge graph automatic generation method and system

    CN119150971A

  • Image search method based on multi-modal algorithm

    CN119226549A

Cited By

  • Multi-mode sound picture storage platform

    CN120639917A

  • Multimodal sound picture storage platform

    CN120639917B

  • Building engineering business opportunity information retrieval agent system and method based on large model

    CN121144503A

  • Big model-based building engineering business opportunity information retrieval agent system and method

    CN121144503B

  • Distributed storage and high-concurrency retrieval system for endoscope video data

    CN122412646A