Visual text coding method and system based on shared semantics and composite external space
Through a visual text encoding method based on shared semantics and composite external space, local features are extracted using multi-layer convolution and bidirectional recurrent networks, and through multiple external space mapping and shared semantic assimilation mechanisms, the problem of inaccurate recognition of cross-modal semantic alignment in multi-category scenarios is solved, achieving efficient and accurate cross-modal retrieval and matching.
Patent Information
- Application Number
- CN202510959561.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing cross-modal semantic alignment methods between visual images and text descriptions have difficulty handling the dynamic changes of multi-category scenes, resulting in inaccurate recognition in scenes with high heterogeneity and subtle differences, and a lack of sufficient attention to key time periods and detailed areas.
A visual text encoding method based on shared semantics and composite external space is adopted. Local features are extracted through multi-layer convolution and bidirectional recurrent networks. Multiple external space mapping and shared semantic assimilation mechanism are utilized, combined with a dynamic attention weighting strategy, to achieve accurate matching between images and texts.
It improves the accuracy and real-time performance of cross-modal semantic alignment, enhances the adaptability to multi-category scenarios, improves the precision of cross-modal retrieval and matching, reduces dependence on specific data distribution, and improves the flexibility and robustness of the system.
Smart Images

Figure CN120632792A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of modal alignment technology, and in particular to a visual text encoding method and system based on shared semantics and composite external space. Background Art
[0002] With the rapid growth of multimedia data and natural language text in various application scenarios, achieving accurate and dynamic cross-modal semantic alignment between visual images and text descriptions has become one of the core challenges in the fields of computer vision and natural language processing. Existing technologies generally use convolutional neural networks (CNNs) to extract image features and recurrent neural networks (RNNs) or Transformers to vectorize and encode text, and then match them based on vector similarity or simple attention mechanisms. However, these traditional methods often have difficulty handling the diversity of image and text content distribution, multi-category scenarios, and application requirements. The matching of local key features and text keywords is insufficient, which can easily lead to inaccurate recognition of scenes with high heterogeneity or subtle differences.
[0003] To improve the effectiveness of cross-modal alignment, some studies have introduced mechanisms such as multi-head attention and feature-level fusion in model structures or training objectives, attempting to capture the local associations between images and text in a more refined manner. However, most of these improvements are limited to a single semantic space or a few categories, and lack the ability to adapt to dynamic changes in large-scale, multi-category environments. Especially in practical applications that require rapid response to multi-category attributes or continuous updates of multimodal data sources, traditional methods still have significant deficiencies in the accuracy and real-time performance of semantic alignment. They are unable to pay sufficient attention to key time periods, detailed areas, or core words in the text, and it is difficult to dynamically adjust the focus of feature fusion according to changes in external conditions.
[0004] To overcome the above difficulties, the present invention proposes a "visual text encoding method based on shared semantics and composite external space" to establish precise and dynamic cross-modal semantic alignment between visual images and text descriptions to achieve efficient and accurate retrieval and matching. Summary of the Invention
[0005] One object of the present invention is to propose a visual text encoding method based on shared semantics and composite external space, which effectively improves the accuracy of cross-modal semantic alignment;
[0006] A visual text encoding method based on shared semantics and composite external space according to an embodiment of the present invention includes the following steps:
[0007] S1. Clean the image and text data, unify the image size, detect the candidate regions, and embed the text characters into indexes to obtain the vectors after mapping.
[0008] S2. Cropping the candidate area to generate a local image block, using multi-layer convolution, nonlinear activation combined with biased weighting to extract local visual features and flatten them to obtain local visual features;
[0009] S3. Use a bidirectional recurrent network to perform gated calculations on the vector obtained after mapping, from left to right and from right to left, respectively, and concatenate the forward and backward hidden states to obtain the local features of the text;
[0010] S4. Define K external spaces, each of which contains visual and text mapping matrices, and map local visual features and local text features to corresponding external spaces respectively;
[0011] S5. Calculate the matching degree and regularization constraints of local visual features and local text features in the external space, and iteratively update the features by sharing semantic assimilation signals;
[0012] S6. Allocate external space-level attention based on the matching score, perform weighted fusion on each external space mapping to obtain a composite external space, and output aggregated features under dimension-level attention;
[0013] S7. Pool or summarize the aggregated features and combine them with contrast loss to complete end-to-end training and generate a cross-modal global representation.
[0014] Furthermore, the S1 includes the following steps:
[0015] S11. The original image data and text data in the cross-modal scenario are recorded as image set I and text set W respectively.
[0016] S12, image dataset I * Any image in is denoted as I(p,q), with height and width H and W respectively; define the target height and target width as H′ and W′ respectively, select the interpolation kernel function K(x,y,p,q), and complete the transformation of each pixel through multi-dimensional weighted summation to obtain the transformed image set
[0017] S13, based on image collection The candidate bounding box set B is defined by the coordinates and width and height of a corner, and the comprehensive scoring function Ψ(b i ) is used to judge b i The effectiveness of the classification includes classification confidence and location accuracy; the optimal solution of the comprehensive scoring function is recorded as Find the optimal bounding box set for each image, filter out the target area with high score and output the accurate bounding box combination, generate the region detection result set R = {R1, R2, ..., R m}, where R i The final detection area corresponding to the i-th image
[0018] S14. Split each text sample in the text dataset W* into several word sequences Mapping tokens to a learnable vector space.
[0019] Furthermore, the S2 includes the following steps:
[0020] S21, according to the partition detection results, each image in the image set Corresponding detection area set R i ={r i1 ,r i2 ,…,r ik}, r ij Represents the coordinates of the upper left corner and width and height information of the jth detection box; for each r ij The coordinates of the image are cropped to generate a local image block Q. ij , obtain several local image blocks, which constitute the input of subsequent local feature encoding;
[0021] S22, introduce a multi-layer convolutional network structure to the local image block Q ij Perform hierarchical convolution and nonlinear activation, and use biased weighted method to obtain output feature maps
[0022] S23, introduce adaptive channel preservation, and convert the last layer feature map F ij L It is considered as a three-dimensional tensor of dimension (C×H×W), where C, H, and W correspond to the number of channels and the height and width of the feature map, respectively, and define the local feature vector The three-dimensional tensor is resized into a one-dimensional vector, and the extracted multi-channel representations are combined by region.
[0023] Furthermore, a weighted method with bias is used to obtain the output feature map as follows:
[0024]
[0025] in Represents the value of the feature map of the previous layer at position (x+p-1, y+q-1) and channel c, c l-1 Indicates the number of channels in the previous layer; W l (p,q,c) is the parameter matrix of the convolution kernel of the lth layer, p and q correspond to the rows and columns of the convolution kernel respectively; b l is the bias term; σ(·) is the nonlinear mapping function; (x, y) represents the current output feature map The coordinate position of the multi-channel feature map sequence {F ij 1 ,Fij 2 ,…,F ij L}.
[0026] Furthermore, the S3 includes the following steps:
[0027] S31, after preliminary text cleaning and word segmentation embedding, we get the text sequence set W*={w1,w2,…,w l}, where each w i is the word vector after embedding, with a dimension of d, and the text sequence set W* is regarded as X=[x1,x2,…,x l ],in Represent the basic vector representation of the i-th word; define a text sequence X with a length of l i(1) Its internal word vector set {x1,x2,…,x l};
[0028] S32, build a forward loop network with multiple gate control mechanisms, and input x at each moment t Perform memory update and state output; traverse t = 1 to t = l, and obtain the forward hidden state set H of the entire sequence (→) ={h1 (→) ,h2 (→) ,…,h1 (→)};
[0029] S33. Under the bidirectional structure, information aggregation is performed on the text sequence from right to left; traversing t = 1 to t = 1, the hidden state set H from right to left can be obtained. (←) ={h1 (←) ,h2 (←) ,…,h l (←)};
[0030] S34, the forward hidden state corresponding to position t and the backward hidden state Do splicing to get bidirectional output h t .
[0031] Furthermore, the S4 includes the following steps:
[0032] S41, let the local visual feature vector set be V = {v1,…,v m}, where each Let the local text feature vector set be T = {t1,…,t n}, where each
[0033] Define K external space matrix groups and in is the mapping matrix of the kth external space in the visual mode, is the mapping matrix of the kth external space in text mode, k = 1, 2, …, K;
[0034] S42, in the definition On this basis, a multi-space mapping relationship for the local visual vector is constructed; let v i Let the i-th visual local feature vector be the mapping result of the i-th visual local feature vector in the k-th external space be recorded as Will As the core parameter of the transformation operation, matrix multiplication is used and explicitly expanded into a cumulative summation form to obtain the mapping result vector Output in the pth dimension;
[0035] Traverse p = 1 to d, and get For all local visual features v i , i = 1…m, after the mapping is completed, the visual transformation result set on the kth external space is obtained
[0036] S43, based on For local features of text t j Perform isomorphic mapping and let the output of the kth external space after the mapping be Similarly, matrix multiplication is performed and expanded explicitly to obtain the mapping result vector Output in the rth dimension;
[0037] Traverse r = 1 to d and get the mapped vector Then the transformed output of all local features of the text is recorded as
[0038] Define the cross observation quantity Z (k) , the mapping results of vision and text in the same external space are accumulated twice to obtain the overall cross-modal interaction strength parameter Z within the same space (k) .
[0039] Furthermore, the S5 includes the following steps:
[0040] S51. For each external space k, obtain a set of visual transformation results. and text transformation result set Define the similarity function The inner product is used in combination with the learnable weight to obtain the matching degree between the i-th visual vector and the j-th text vector in the external space k; traversing i=1…m, j=1…n, the matching score matrix of all visual and text features in the external space is obtained.
[0041] S52, let L r is the joint regularization term for cross-modal matching; let r k Represents the regular weight corresponding to the external space k, and defines the total objective function L total Represents the sum of all external space matches and regular expressions:
[0042]
[0043] Where Ψ(·) is a function that measures the difference between the visual and textual matching scores, used to highlight the distinction between high similarity and low similarity; is the regularization term inside the external space k, which is used to constrain the distribution of visual and text features to form a consistent trend; r k represents the regularization weight of the kth external space, balancing the matching degree and the penalty term;
[0044] S53. Obtain cross-modal assimilation signals through gradient backtracking or other forms of optimization algorithms and Indicates the correction trend of visual and text features in the kth external space;
[0045] By and Feedback to each vector Internally, vision and text are forced to keep synchronized and adjusted in this external space, forming a shared semantic assimilation.
[0046] Furthermore, the step S6 includes the following steps:
[0047] S61. In the external space k, the overall matching score of the input local feature f is R (f,k) ;
[0048] The matching scores of all spaces are normalized by the normalized activation function to obtain the attention weight α (k) , α (k) is the attention coefficient corresponding to the k-th space, satisfying 0<α (k) <1 and
[0049] S62, perform weighted summation on the multi-space mapping results to obtain the composite external space M*, and convert the mapping matrix corresponding to each space k into Combined through attention coefficient;
[0050] By weighting the elements of K spaces, a composite external space M is formed that comprehensively considers multi-semantic preferences and dynamic matching. comp ;
[0051] S63, composite external space M comp Introducing dimension-level attention aggregation strategy;
[0052] According to M comp The composite mapping defines the eigenvector μi obtained after the mapping, which is expressed using matrix multiplication and introduces an additional dimension attention parameter μ i (r), r = 1, 2, ..., d, μi(r) is the aggregate output value at dimension r;
[0053] Traversing r=1 to d can obtain the complete vector The output combination corresponding to all input features is M = {μ1,μ2,…,μn}.
[0054] Furthermore, the step S7 includes the following steps:
[0055] S71. Collect local features of each image or text, and aggregate or pool each dimension in the composite external space to obtain a global representation vector of uniform length;
[0056] S72, feeding the global representation vector and its corresponding cross-modal matching relationship into an adaptive contrast loss, so that the distance between similar sample vectors is reduced and the distance between dissimilar samples is increased, thereby enhancing cross-modal discrimination;
[0057] S73, based on the cross-modal matching score and its assimilation signal, that is, the spatial level attention coefficient {α k} and dimension-level attention coefficient {β r}, by differential weight assignment, correct the visual-text alignment distance in the global representation vector space;
[0058] S74. Combine the global representation vector with the matching loss, cooperate with the shared semantic constraints, perform end-to-end network training and output the final cross-modal global representation.
[0059] The present invention also provides a visual text encoding system based on cross-modal shared semantics and composite external space, comprising a data processing module, a local visual feature extraction module, a text local feature extraction module, a mapping module, a feature iterative update module, a fusion module, and a generation module;
[0060] The data processing module is used to clean the image and text data, unify the image size, detect candidate regions, and index and embed text characters;
[0061] The local visual feature extraction module is used to crop the candidate area to generate local image blocks, and extract local visual features and flatten them using multi-layer convolution, nonlinear activation and biased weighting to obtain local visual features;
[0062] The text local feature extraction module is used to perform gated calculations on the text sequence from left to right and from right to left using a bidirectional recurrent network, and concatenate the forward and backward hidden states to obtain the text local features;
[0063] The mapping module is used to define K external spaces, each of which contains visual and text mapping matrices, mapping local visual features and local text features to the corresponding external space respectively;
[0064] The feature iterative update module is used to calculate the matching degree and regularization constraints of local visual features and local text features in the external space, and iteratively update the features by sharing semantic assimilation signals;
[0065] The fusion module is used to allocate external space-level attention based on the matching score, perform weighted fusion of each external space mapping to obtain a composite external space, and output aggregated features under dimension-level attention;
[0066] The generation module is used to pool or summarize the aggregated features, combine the contrast loss to complete end-to-end training, and generate cross-modal global representations.
[0067] Compared with the prior art, the present invention has at least the following beneficial effects:
[0068] (1) The present invention adopts multiple external space mapping and shared semantic assimilation mechanisms, which not only enhances the capture of key features in the visual and text encoding process, but also achieves fine distinction of multi-dimensional differences between images and text through parallel mapping and semantic calibration, enabling the model to maintain stable alignment effect and strong adaptability in highly heterogeneous and multi-category scenarios;
[0069] (2) The present invention introduces a dynamic attention weighting strategy in the multimodal fusion process, hierarchically aggregates the matching degree between visual features and text semantics in different external spaces, and assigns attention weights based on real-time scoring, thereby accurately highlighting the correlation between key image areas and text roots, effectively improving the accuracy of cross-modal retrieval and matching, and enhancing the ability to recognize subtle differences and local features;
[0070] (3) The invention utilizes a design concept that combines multi-level feature abstraction with regularization to dynamically adjust the feature fusion depth for different categories of attributes, significantly reducing the reliance on specific data distributions or domain priors, and reducing the burden of repeated parameter adjustments or excessive labeling in large-scale practical applications, enabling the system to maintain high flexibility and robustness when facing diverse graphic and text data; BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0072] Figure 1 This is a flowchart of a visual text encoding method based on shared semantics and composite external space proposed by the present invention;
[0073] Figure 2 A flowchart for constructing multi-external space mapping for a visual text encoding method based on shared semantics and composite external space proposed by the present invention; DETAILED DESCRIPTION
[0074] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner, and therefore only show components relevant to the present invention.
[0075] refer to Figure 1 , a visual text encoding method based on shared semantics and composite external space, comprising the following steps:
[0076] S1, cleaning the image and text data, unifying the image size, detecting candidate areas and indexing and embedding the text characters; S1 includes the following steps:
[0077] S11, the original image data and text data in the cross-modal scenario are respectively recorded as image sets I = {I1, I2, ..., I m} and text set W={w1,w2,…,w n}, where m represents the number of images and n represents the number of texts; remove abnormal resolution images and remove image elements with severe noise to obtain the cleaned image dataset I * ; At the same time, the text data is reviewed for character integrity and encoded, and blank placeholders and meaningless symbols are removed to obtain the cleaned text dataset W * ;
[0078] S12, image data set I obtained according to step S11 * , let any image be denoted as I(p,q), with height and width as H and W respectively; to ensure the size consistency of the subsequent feature extraction network, define the target height and target width as H′ and W′ respectively; select the interpolation kernel function K(x,y,p,q), and complete the transformation of each pixel through multi-dimensional weighted summation, and record the resulting image as
[0079]
[0080] Where x = 1, 2, ..., H′, y = 1, 2, ..., W′, I(p, q) represents the pixel value at position (p, q) of the original image, and K(x, y, p, q) represents the interpolation weight from (p, q) to (x, y);
[0081] Through the interpolation process of the above formula, all images can be unified to the same size and retain detail information to obtain the transformed image set
[0082] S13, obtained in step S12 On this basis, in order to extract the target area in the image, a candidate bounding box set B = {b1, b2, ..., b k}, where b i =(x i ,y i ,w i ,h i ) are the coordinates of the upper left corner and the width and height of the i-th candidate box;
[0083] Introducing the comprehensive scoring function Ψ(b i ) is used to judge b i The effectiveness of the classification includes classification confidence and location accuracy; the optimal solution of the comprehensive scoring function is recorded as To obtain high-confidence detection results, the following formula is used to find the optimal bounding box set for each image:
[0084]
[0085] Among them, φ cls represents the classification confidence, φ loc represents the positioning accuracy, α and β are weight coefficients used to balance the classification and positioning targets;
[0086] After solving the above equation, the target areas with high scores are screened out and the precise bounding box combination is output to generate the region detection result set R = {R1, R2, ..., R m}, where R i The final detection area corresponding to the i-th image;
[0087] S14. In the text dataset W* obtained in step S11, each text sample can be split into several word sequences To map word units to a learnable vector space, define the embedding matrix Where d represents the embedding dimension and V represents the dictionary size; according to the index encoding function
[0088]
[0089] in, is the text word w nj Vector representation of ;
[0090] S2, cropping the candidate region to generate a local image block, extracting local visual features using multi-layer convolution and nonlinear activation, and flattening the local visual features; S2 includes the following steps:
[0091] S21, according to the partition detection result of step S1, the image set is recorded as Each image Corresponding detection area set R i ={r i1 ,r i2 ,…,r ik}, r ij Represents the coordinates of the upper left corner and width and height information of the jth detection box; for each r ij The coordinates of the image are cropped to generate a local image block Q. ij ; This sub-step outputs several local image blocks, which constitute the input for subsequent local feature encoding;
[0092] S22, introduce a multi-layer convolutional network structure to the local image block Q ij Perform hierarchical convolution and nonlinear activation, record the number of network layers as L, the convolution kernel size as (u×v), and the number of output channels as c l ; Define the convolution operation H l (·) is the convolution function of the first layer, and the output feature map is obtained by using a weighted method with bias
[0093]
[0094] in Represents the value of the feature map of the previous layer at position (x+p-1, y+q-1) and channel c, c l-1 Indicates the number of channels in the previous layer; W l (p,q,c) is the parameter matrix of the convolution kernel of the lth layer, p and q correspond to the rows and columns of the convolution kernel respectively; b l is the bias term; σ(·) is the nonlinear mapping function; (x, y) represents the current output feature map The coordinate position of
[0095] Through the above formula, under the action of multi-layer stacked convolution and activation, each local image block Q ij The high-level semantic features of the ij 1 ,F ij 2 ,…,F ij L};
[0096] S23, in order to highlight the multi-dimensional information of the local area and avoid the loss of details caused by global average pooling, an adaptive channel preservation strategy is introduced; the last layer feature map F ij L Treated as a three-dimensional tensor of dimension (C×H×W), where C, H, and W correspond to the number of channels and the height and width of the feature map respectively; define the local feature vector The three-dimensional tensor is rearranged into a one-dimensional vector by the following formula to fully preserve the local information:
[0097]
[0098] Where R(·) means flattening the three-dimensional tensor in channel priority order; v i,j Maintain C×H×W elements to record fine-grained information of the local area;
[0099] The extracted multi-channel representations are combined into V by region i ={v i1 ,v i2 ,…,v ik};
[0100] S3. Use a bidirectional recurrent network to perform gated calculations on the text sequence from left to right and from right to left, respectively, and concatenate the forward and backward hidden states to obtain local text features. This specifically includes the following steps:
[0101] S31, after completing the preliminary text cleaning and word segmentation embedding in step S1, the text sequence set W*={w1,w2,…,w l}, where each w i is the word vector after embedding, with a dimension of d; the text sequence set is regarded as X = [x1, x2, ..., x l ],in Represent the basic vector representation of the i-th word; to ensure the input sequence of the bidirectional recurrent network, define a text sequence X with a length of l i(1) Its internal word vector set {x1,x2,…,x l};
[0102] S32, in order to obtain the contextual features of the text sequence set in the order dimension, a forward loop network with a multi-gating mechanism is constructed, and the input x at each moment is t Perform memory update and state output; let the hidden state of the forward network be recorded as The input gate, forget gate, and output gate are The cell state is Select the weight matrix Bias term and nonlinear mapping function σ(·), and define the Hadamard product as ⊙;
[0103] The gate calculation relationship of the forward network:
[0104]
[0105] In the above formula, x t Represents the word vector at the current moment, represents the hidden state at the previous moment, Represents the cell state at the previous moment; through this recursive structure, the semantic dependency from left to right is captured; traverse t = 1 to t = l to obtain the forward hidden state set H of the entire sequence (→) ={h1 (→) ,h2 (→) ,…,h l (→)};
[0106] S33. Under the bidirectional structure, it is also necessary to aggregate the information of the text sequence set from right to left; let the hidden state of the backward network be Let the input gate, forget gate, and output gate be The cell state is Weight Matrix Bias term Relatively independent of the forward network; for the word vector x indexed from back to front l-t+1 , the gating mechanism of the backward network is shown as follows:
[0107]
[0108] Traversing t=1 to t=l, we can get the hidden state set H from right to left. (←) ={h1 (←) ,h2 (←) ,…,h l (←)};
[0109] S34, the forward hidden state corresponding to position t and the backward hidden state Do splicing to get bidirectional output h t ; Define the fusion function G(·) as:
[0110]
[0111] Where [·;·] is a vector concatenation operation; traversing from t = 1 to t = l, the complete text local feature set H = {h1, h2, ..., h1} can be obtained. Each h t Contains bidirectional semantic information from left to right and from right to left;
[0112] S4. Define K external spaces, each containing visual and text mapping matrices, and map local visual features and text features to the corresponding external space respectively; refer to Figure 2 , including the following steps:
[0113] S41, let the local visual feature vector set be V = {v1,…,v m}, where each Let the local text feature vector set be T = {t1,…,t n}, where each
[0114] In order to distinguish multiple semantic preferences such as scenes, actions, entities, etc. in cross-modal scenarios, K external space matrix groups are defined and in is the mapping matrix of the kth external space in the visual mode, is the mapping matrix of the kth external space in text mode, k = 1, 2, …, K;
[0115] S42, defined in sub-step S41 On this basis, a multi-space mapping relationship for the local visual vector is constructed; let v i Let the i-th visual local feature vector be its mapping result in the k-th external space be recorded as Will As the core parameter of the transformation operation, matrix multiplication is used and explicitly expanded into the cumulative sum form:
[0116]
[0117] Among them, v i (q) represents the local visual feature v i The value in the qth dimension; Represents the weight of the k-th visual external space mapping matrix at row p and column q; Represents the mapping result vector Output in the pth dimension;
[0118] Traversing from p=1 to d, we can get For all local visual features v i , i = 1…m, after the mapping is completed, the visual transformation result set on the kth external space can be obtained
[0119] S43, based on sub-step S41 For local features of text t j Perform isomorphic mapping and let the output of the kth external space after the mapping be Similarly, we use matrix multiplication and make an explicit expansion:
[0120]
[0121] where t j (s) represents the local feature of the text t j The value in the sth dimension; Represents the weight of the k-th text external space mapping matrix at row r and column s; Represents the mapping result vector Output in the rth dimension;
[0122] Traversing r=1 to d, we can get the mapped vector Then the transformed output of all local features of the text is recorded as
[0123] In order to further align the visual and text modalities across space in the subsequent stage, the cross-observation Z is defined (k) , perform secondary accumulation on the mapping results of vision and text in the same external space; let is the weight for measuring the correlation between the i-th visual feature and the j-th text feature in the k-th external space:
[0124]
[0125] in Represents the inner product of the vectors of vision and text after being transformed in the same external space; To control the coefficient of the matching evaluation between the two in this space; Z (k) Quantified the overall cross-modal interaction strength within the same space;
[0126] S5. Calculate the matching degree and regular constraints between vision and text in each space, and iteratively update the features by sharing semantic assimilation signals; transform the local visual features {v i} and local text features {t j} is mapped to the kth external space, and then based on the inner product equation Calculate the matching degree, add the regularization term, and reversely optimize to obtain the shared assimilation signal Δv i ,Δt j The shared assimilation signal Δv i ,Δt j For iterative update {v i},{t j}, specifically including the following steps:
[0127] S51. In step S4, for each external space k, obtain a set of visual transformation results and text transformation result set
[0128]
[0129] in, Represents the vector of the local visual feature in the kth external space; Represents the vector of the local feature of the text in the kth external space; α (k) is a learnable scalar parameter used to adjust the numerical scale of different spaces; Represents the similarity benchmark generated by the inner product operation of the two; traversing i = 1…m, j = 1…n, the matching score matrix of all visual and text features in this external space can be obtained
[0130] S52. After completing the matching calculation, it is necessary to balance the joint alignment degree in each external space, so that L r is a joint regularization term for cross-modal matching, which is used to establish a semantic assimilation effect between visual and text features; let r k Represents the regular weight corresponding to the external space k, and defines the total objective function L total , which represents the sum of all external space matches and regular expressions:
[0131]
[0132] Where Ψ(·) is a function that measures the difference between the visual and textual matching scores, used to highlight the distinction between high similarity and low similarity; is the regularization term inside the external space k, which is used to constrain the distribution of visual and text features to form a consistent trend; r k represents the regularization weight of the kth external space, balancing the matching degree and the penalty term;
[0133] S53, in step S52, the total objective function L total Based on this, we can obtain cross-modal assimilation signals through gradient backtracking or other forms of optimization algorithms. and The cross-modal assimilation signal represents the correction trend of visual features and text features in the kth external space; the shared semantic assimilation operation is defined as:
[0134]
[0135] in, is the step size coefficient, which is used to determine the update amplitude on the visual and text sides; and Represent the negative gradient of the total objective function with respect to the visual and text vectors respectively; and Feedback to each vector Inside, the visual and text are forced to keep synchronized and adjusted in this external space, forming a shared semantic assimilation;
[0136] S6. Allocate external space-level attention based on the matching score, perform weighted fusion on each external space mapping, and output aggregated features under dimension-level attention; including the following steps:
[0137] S61, under the shared semantic constraint of step S5, the visual end and the text end matching measure in K external spaces has been obtained; record the overall matching score of the input local feature f in the external space k as R (f,k) ;
[0138] The matching scores of all spaces are normalized by the normalized activation function to obtain the attention weight α (k) :
[0139]
[0140] Among them, R (f,k) represents the matching score of feature f in the kth external space, β is a learnable or settable amplification coefficient used to enhance the difference, and competitively normalizes the matching degree of each space through the exponential operation in the numerator and denominator; α (k) is the attention coefficient corresponding to the k-th space, satisfying 0<α (k) <1 and
[0141] S62, α obtained in step S61 (k) On this basis, the multi-space mapping results of step S4 are weighted summed to obtain the composite external space M*; the mapping matrix corresponding to each space k is Combined by attention coefficient; let the mapping matrix Represents the mapping function parameter matrix of the k-th space, then the construction method of the composite external space is:
[0142]
[0143] Among them, d is the dimension of the feature vector, is the coefficient on row p and column q of the k-th space mapping parameter;
[0144] α (k) represents the attention weight calculated in step S61; by weighting the elements of the K spaces, a composite external space M is formed that comprehensively considers multi-semantic preferences and dynamic matching. comp ;
[0145] S63, the composite external space M obtained in sub-step S62 compIn order to further highlight the key information of local features, a dimension-level attention aggregation strategy is introduced; the input feature vector set is Γ={γ1,γ2,…,γn}, each Indicates the local visual or text features that need to be aggregated;
[0146] According to M comp The composite mapping of , defines the eigenvector μi obtained after mapping, using matrix multiplication to express and additionally introduce dimension attention parameters:
[0147]
[0148] Among them, M comp (r,q) is the mapping parameter of the composite external space at row r and column q, γ i (q) is the value of the local feature vector in the qth dimension; ω(r) is the attention bias coefficient specifically for the rth dimension, represents nonlinear activation; μi(r) is the aggregated output value at dimension r;
[0149] Traversing r=1 to d can obtain the complete vector Then record the output combination corresponding to all input features as M = {μ1,μ2,…,μn};
[0150] S7. Pool or aggregate the aggregated features and combine them with contrast loss to complete end-to-end training and generate the final cross-modal global representation. This includes the following steps:
[0151] S71. Collect local features of each image or text based on the composite external space generated in step S6, and aggregate or pool each dimension in this space to obtain a high-dimensional representation vector of uniform length, providing a feature basis for downstream cross-modal tasks.
[0152] S72, feeding the global representation vector and its corresponding cross-modal matching relationship into an adaptive contrast loss, so that the distance between similar sample vectors is reduced and the distance between dissimilar samples is increased, thereby enhancing cross-modal discrimination;
[0153] S73, based on the cross-modal matching score and its assimilation signal, that is, the spatial level attention coefficient {α k} and dimension-level attention coefficient {β r By assigning differentiated weights, the visual-text alignment distance in the global representation vector space is corrected, ensuring that difficult samples achieve better recognition and matching performance in this space.
[0154] S74. Combine the global representation vector with the matching loss, cooperate with the shared semantic constraints generated in the previous steps, perform end-to-end network training and output the final cross-modal global representation, providing high-reliability vector support for retrieval and matching that adapts to multiple scenarios and multiple semantics.
[0155] Example 1:
[0156] A large e-commerce platform in City A received hundreds of user-posted product images and text. To improve listing review efficiency and reduce the risk of false advertising, the platform deployed the method described in this invention to automatically detect and evaluate the matching between each product image and text description.
[0157] First, the system cleans and resizes all images, removing images with severely abnormal resolutions and filtering out obviously distorted text descriptions. Then, using candidate bounding box detection technology, it automatically searches for key visual areas (trademarks, characteristic patterns, and text labels) within the product images, cropping the located local areas to generate several image blocks. On the visual side, the system extracts high-level features through a multi-layer convolutional network and flattens them, ensuring that visual vectors of consistent length are output for product images of different sizes.
[0158] On the text side, the platform segmented product titles and key descriptions and mapped them into word vector space. It then used a bidirectional recurrent network to capture semantic dependencies from left to right and right to left, concatenating them into bidirectional hidden states and ultimately outputting a text feature sequence.
[0159] To achieve cross-modal alignment, the system maps visual and textual features across multiple external spaces and calculates the matching score of the inner product and regularization constraints. During the training phase, the visual and textual distributions are iteratively modified by sharing semantic assimilation signals. After completing the multi-space mapping, the system allocates attention to the matching scores of each space, performs a weighted fusion of the mapping matrices to form a composite external space, and applies attention aggregation at the dimensional level to highlight key features. The final high-dimensional representation vector is optimized using a contrastive loss function to ensure matching accuracy and robustness.
[0160] In actual operation, the system detected the product information of "XX leather shoes" uploaded by the user; because the light was dim when the picture was taken, traditional detection schemes could not accurately identify the texture of the shoe upper, which made it easy to confuse "leather shoes" with "sports shoes"; however, the present invention obtained key image blocks (such as shoe heads and trademarks) through local detection, combined with the context of high-frequency words such as "genuine leather" and "business" after text-end word embedding; after multi-external space mapping, the cross-modal matching degree between the two was significantly improved, and the final score obtained by integrating the attention mechanism was 0.81. The system determined that the product description was "highly authentic".
[0161] The visual-text encoding method of the present invention effectively improves the efficiency and accuracy of image-text consistency detection in e-commerce scenarios, fully demonstrating the excellent performance of cross-modal shared semantics and dynamic fusion of external space in practical applications.
[0162] Example 2: Based on the concept of the method, the present invention can also provide a visual text encoding system based on cross-modal shared semantics and composite external space, including a data processing module, a local visual feature extraction module, a text local feature extraction module, a mapping module, a feature iterative update module, a fusion module, and a generation module;
[0163] The data processing module is used to clean the image and text data, unify the image size, detect the candidate area, and index and embed the text characters to obtain the vector after mapping;
[0164] The local visual feature extraction module is used to crop the candidate area to generate local image blocks, and extract local visual features and flatten them using multi-layer convolution, nonlinear activation and biased weighting to obtain local visual features;
[0165] The text local feature extraction module is used to perform gated calculations on the vector obtained after mapping using a bidirectional recurrent network from left to right and from right to left, and then concatenate the forward and backward hidden states to obtain the text local features.
[0166] The mapping module is used to define K external spaces, each of which contains visual and text mapping matrices, mapping local visual features and local text features to the corresponding external space respectively;
[0167] The feature iterative update module is used to calculate the matching degree and regularization constraints of local visual features and local text features in the external space, and iteratively update the features by sharing semantic assimilation signals;
[0168] The fusion module is used to allocate external space-level attention based on the matching score, perform weighted fusion of each external space mapping to obtain a composite external space, and output aggregated features under dimension-level attention;
[0169] The generation module is used to pool or summarize the aggregated features, combine the contrast loss to complete end-to-end training, and generate cross-modal global representations.
[0170] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A visual text encoding method based on shared semantics and composite external space, characterized in that: The steps include: S1. Clean the image and text data, unify the image size, detect the candidate regions, and embed the text characters into indexes to obtain the vectors after mapping. S2. Cropping the candidate area to generate a local image block, using multi-layer convolution, nonlinear activation combined with biased weighting to extract local visual features and flatten them to obtain local visual features; S3. Use a bidirectional recurrent network to perform gated calculations on the vector obtained after mapping, from left to right and from right to left, respectively, and concatenate the forward and backward hidden states to obtain the local features of the text; S4. Define K external spaces, each of which contains visual and text mapping matrices, and map local visual features and local text features to corresponding external spaces respectively; S5. Calculate the matching degree and regularization constraints of local visual features and local text features in the external space, and iteratively update the features by sharing semantic assimilation signals; S6. Allocate external space-level attention based on the matching score, perform weighted fusion on each external space mapping to obtain a composite external space, and output aggregated features under dimension-level attention; S7. Pool or summarize the aggregated features and combine them with contrast loss to complete end-to-end training and generate a cross-modal global representation.
2. A visual text encoding method based on shared semantics and composite external space according to claim 1, characterized in that: Said S1 comprises the following steps: S11. The original image data and text data in the cross-modal scenario are recorded as image set I and text set W respectively. Abnormal resolution images are removed and image elements containing severe noise are eliminated to obtain the cleaned image dataset I. * ; Perform character integrity review and encoding filtering on the text data, remove blank placeholders and meaningless symbols, and obtain the cleaned text dataset W * ; S12, image dataset I * Any image in is denoted as I(p,q), with height and width H and W respectively; define the target height and target width as H′ and W′ respectively, select the interpolation kernel function K(x,y,p,q), and complete the transformation of each pixel through multi-dimensional weighted summation to obtain the transformed image set S13, based on image collection The candidate bounding box set B is defined by the coordinates and width and height of a corner, and the comprehensive scoring function Ψ(b i ) is used to judge b i The effectiveness of the classification includes classification confidence and location accuracy; the optimal solution of the comprehensive scoring function is recorded as Find the optimal bounding box set for each image, filter out the target area with high score and output the accurate bounding box combination, generate the region detection result set R = {R1, R2, ..., R m }, where R i The final detection area corresponding to the i-th image S14. Split each text sample in the text dataset W* into several word sequences Mapping tokens to a learnable vector space.
3. The method for encoding visual text based on shared semantics and composite external space according to claim 1, characterized in that: The S2 comprises the following steps: S21, according to the partition detection results, each image in the image set Corresponding detection area set R i ={r i1 ,r i2 ,…,r ik }, r ij Represents the coordinates of the upper left corner and width and height information of the jth detection box; for each r ij The coordinates of the image are cropped to generate a local image block Q. ij , obtain several local image blocks, which constitute the input of subsequent local feature encoding; S22, introduce a multi-layer convolutional network structure to the local image block Q ij Perform hierarchical convolution and nonlinear activation, and use biased weighted method to obtain output feature maps S23, introduce adaptive channel preservation, and convert the last layer feature map F ij L It is considered as a three-dimensional tensor of dimension (C×H×W), where C, H, and W correspond to the number of channels and the height and width of the feature map, respectively, and define the local feature vector The three-dimensional tensor is resized into a one-dimensional vector, and the extracted multi-channel representations are combined by region.
4. The method for encoding visual text based on shared semantics and composite external space according to claim 3, characterized in that: Use biased weighted method to obtain output feature map as follows: in Represents the value of the feature map of the previous layer at position (x+p-1, y+q-1) and channel c, c l-1 Indicates the number of channels in the previous layer; W l (p,q,c) is the parameter matrix of the convolution kernel of the lth layer, p and q correspond to the rows and columns of the convolution kernel respectively; b l is the bias term; σ(·) is the nonlinear mapping function; (x, y) represents the current output feature map The coordinate position of the multi-channel feature map sequence {F ij 1 ,F ij 2 ,…,F ij L }.
5. The method for encoding visual text based on shared semantics and composite external space according to claim 1, characterized in that: The S3 includes the following steps: S31, after preliminary text cleaning and word segmentation embedding, we get the text sequence set W*={w1,w2,…,w l }, where each w i is the word vector after embedding, with a dimension of d, and the text sequence set W* is regarded as X=[x1,x2,…,x l ],in Represent the basic vector representation of the i-th word; define a text sequence X with a length of l i(1) Its internal word vector set {x1,x2,…,x l }; S32, build a forward loop network with multiple gate control mechanisms, and input x at each moment t Perform memory update and state output; traverse t = 1 to t = l, and obtain the forward hidden state set H of the entire sequence (→) ={h1 (→) ,h2 (→) ,…,h l (→) }; S33. Under the bidirectional structure, information aggregation is performed on the text sequence set from right to left; traversing t = 1 to t = 1, the hidden state set H from right to left can be obtained. (←) ={h1 (←) ,h2 (←) ,…,h l (←) }; S34, the forward hidden state corresponding to position t and the backward hidden state Do splicing to get bidirectional output h t .
6. The method for encoding visual text based on shared semantics and composite external space according to claim 1, characterized in that: The S4 comprises the following steps: S41, let the local visual feature vector set be V = {v1,…,v m }, where each Let the local text feature vector set be T = {t1,…,t n }, where each Define K external space matrix groups and in is the mapping matrix of the kth external space in the visual mode, is the mapping matrix of the kth external space in text mode, k = 1, 2, …, K; S42, in the definition On this basis, a multi-space mapping relationship for the local visual vector is constructed; let v i Let the i-th visual local feature vector be the mapping result of the i-th visual local feature vector in the k-th external space be recorded as Will As the core parameter of the transformation operation, matrix multiplication is used and explicitly expanded into a cumulative summation form to obtain the mapping result vector Output in the pth dimension; Traverse p = 1 to d, and get For all local visual features v i , i = 1…m, after the mapping is completed, the visual transformation result set on the kth external space is obtained S43, based on For local features of text t j Perform isomorphic mapping and let the output of the kth external space after the mapping be Similarly, matrix multiplication is performed and expanded explicitly to obtain the mapping result vector Output in the rth dimension; Traverse r = 1 to d and get the mapped vector Then the transformed output of all local features of the text is recorded as Define the cross observation quantity Z (k) , the mapping results of vision and text in the same external space are accumulated twice to obtain the overall cross-modal interaction strength parameter Z within the same space (k) .
7. The method for encoding visual text based on shared semantics and composite external space according to claim 1, characterized in that: The S5 comprises the following steps: S51. For each external space k, obtain a set of visual transformation results and a set of text transformation results respectively; define a similarity function The inner product is used in combination with the learnable weight to obtain the matching degree between the i-th visual vector and the j-th text vector in the external space k; traversing i=1…m, j=1…n, the matching score matrix of all visual and text features in the external space is obtained. S52, let L r is the joint regularization term for cross-modal matching, r k Represents the regular weight corresponding to the external space k, and defines the total objective function L total Represents the sum of all external space matches and regular expressions: Where Ψ(·) is a function that measures the difference between the visual and textual matching scores, used to highlight the distinction between high similarity and low similarity; is the regularization term inside the external space k, which is used to constrain the distribution of visual and text features to form a consistent trend; r k represents the regularization weight of the kth external space, balancing the matching degree and the penalty term; S53. Obtain cross-modal assimilation signals through gradient backtracking or other forms of optimization algorithms and Indicates the correction trend of visual and text features in the kth external space; By and Feedback to each vector Internally, vision and text are forced to keep synchronized and adjusted in this external space, forming a shared semantic assimilation.
8. The method for encoding visual text based on shared semantics and composite external space according to claim 1, characterized in that: The S6 comprises the following steps: S61. In the external space k, the overall matching score of the input local feature f is R (f,k) ; The matching scores of all spaces are normalized by the normalized activation function to obtain the attention weight α (k) , α (k) is the attention coefficient corresponding to the k-th space, satisfying 0<α (k) <1 and S62, perform weighted summation on the multi-space mapping results to obtain the composite external space M*, and convert the mapping matrix corresponding to each space k into Combined through attention coefficient; By weighting the elements of K spaces, a composite external space M is formed that comprehensively considers multi-semantic preferences and dynamic matching. comp ; S63, composite external space M comp Introducing dimension-level attention aggregation strategy; According to M comp The composite mapping defines the eigenvector μi obtained after the mapping, which is expressed using matrix multiplication and introduces an additional dimension attention parameter μ i (r), r = 1, 2, ..., d, μi(r) is the aggregate output value at dimension r; Traversing r=1 to d can obtain the complete vector The output combination corresponding to all input features is M = {μ1,μ2,…,μn}.
9. The method for encoding visual text based on shared semantics and composite external space according to claim 1, characterized in that: The S7 comprises the following steps: S71. Collect local features of each image or text, and aggregate or pool each dimension in the composite external space to obtain a global representation vector of uniform length; S72, feeding the global representation vector and its corresponding cross-modal matching relationship into an adaptive contrast loss, so that the distance between similar sample vectors is reduced and the distance between dissimilar samples is increased, thereby enhancing cross-modal discrimination; S73, based on the cross-modal matching score and its assimilation signal, that is, the spatial level attention coefficient {α k } and dimension-level attention coefficient {β r }, by differential weight assignment, correct the visual-text alignment distance in the global representation vector space; S74. Combine the global representation vector with the matching loss, cooperate with the shared semantic constraints, perform end-to-end network training and output the final cross-modal global representation.
10. A visual text encoding system based on cross-modal shared semantics and composite external space, characterized by: It includes data processing module, local visual feature extraction module, text local feature extraction module, mapping module, feature iterative update module, fusion module and generation module; The data processing module is used to clean the image and text data, unify the image size, detect the candidate area, and index and embed the text characters to obtain the vector after mapping; The local visual feature extraction module is used to crop the candidate area to generate local image blocks, and extract local visual features and flatten them using multi-layer convolution, nonlinear activation and biased weighting to obtain local visual features; The text local feature extraction module is used to perform gated calculations on the vector obtained after mapping using a bidirectional recurrent network from left to right and from right to left, and then concatenate the forward and backward hidden states to obtain the text local features. The mapping module is used to define K external spaces, each of which contains visual and text mapping matrices, mapping local visual features and local text features to the corresponding external space respectively; The feature iterative update module is used to calculate the matching degree and regularization constraints of local visual features and local text features in the external space, and iteratively update the features by sharing semantic assimilation signals; The fusion module is used to allocate external space-level attention based on the matching score, perform weighted fusion of each external space mapping to obtain a composite external space, and output aggregated features under dimension-level attention; The generation module is used to pool or summarize the aggregated features, combine the contrast loss to complete end-to-end training, and generate cross-modal global representations.
Citation Information
Cited By
Visual token generation method and device based on shared index, equipment and medium
CN120953425A
Method, apparatus and medium for generating visual token based on shared index
CN120953425B
Controllable three-dimensional graphic content generation method and system based on multi-image fusion
CN121095470A
Multi-scale visual positioning method and system based on semantic consistency guidance
CN121353998A
A multi-scale visual positioning method and system based on semantic consistency guidance
CN121353998B