Method for optimizing image recognition and action tracking system of touch interactive terminal for literature and blog

Through multi-modal image feature extraction and multi-feature matching mechanism, the problem of insufficient background matching accuracy of touch interactive terminals in cultural and museum venues is solved, and higher background matching accuracy and user interaction experience are improved.

CN120182773AInactive Publication Date: 2025-06-20GUANG DONG ULTRA PICTURES CULTURE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510665857.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the touch interactive terminals of cultural and museum venues, the background matching accuracy of image recognition and action tracking systems is insufficient, resulting in poor user interaction experience.

Method used

Multimodal images (color, depth, infrared) are used for feature extraction, and weights are generated through local quality evaluation, combined with feature pyramid fusion and multi-head LSTM analysis, user vectors and pose labels are generated, and the optimal background is finally selected through the multi-feature matching mechanism.

Benefits of technology

It significantly improves background matching accuracy and interactive immersion, improves the stability and responsiveness of the user experience, and enhances the continuity and fluency of visual presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182773A_ABST
    Figure CN120182773A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to an optimization method of an image recognition and action tracking system of a touch interactive terminal for literature and blog. Firstly, a multi-modal image is obtained, features of the multi-modal image are extracted to obtain a first image, and then the first image is weighted based on the weight of local quality evaluation to obtain a second image. And fusing the second image features by using the feature pyramid and decoding to obtain a third image, and inputting the third image into a segmentation head and a key point detection head to obtain a segmentation image and key point coordinates. Next, key point changes are calculated to obtain a first sequence; and carrying out weighted fusion on the third image and the first sequence, and carrying out multi-head LSTM analysis to obtain a first vector, a second vector and an attitude label. And finally, based on the first vector, the second vector and the attitude label, selecting an optimal background through a multi-feature matching mechanism. According to the method, the background matching precision and the interaction immersion are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to an optimization method for an image recognition and motion tracking system of a touch interaction terminal for cultural relics and museums. Background Art

[0002] With the increasing attention of users to cultural relics and museums, users have put forward higher requirements for the functions of touch interaction devices. The prior art provides some solutions: one method is to provide a static image library for users to select, which depends on manual switching by users and affects the user experience. Another type of automatic matching method extracts feature vectors of user actions and postures, and uses a pattern matching method based on simple similarity calculation and threshold discrimination to select the background. The quality of the feature vectors of user actions and postures extracted by these methods and the robustness of the matching method need to be further improved, and there are often problems of incorrect matching between the selected background and the actual state of the user. When the matching is incorrect, it will affect the visual coherence of the subsequent synthesis of the user and the background image, and affect the user's interaction experience.

[0003] Therefore, an optimization method for an image recognition and motion tracking system of a touch interaction terminal for cultural relics and museums is proposed. Summary of the Invention

[0004] The optimization method for the image recognition and motion tracking system of the touch interaction terminal for cultural relics and museums of the present invention significantly improves the background matching accuracy and interaction immersion. First, a multi-modal image is obtained, and the multi-modal image features are extracted to obtain a first image. Then, the first image is weighted based on the weights of local quality evaluation to obtain a second image. The second image features are fused by a feature pyramid and decoded to obtain a third image, and the third image is input into a segmentation head and a key point detection head to obtain a segmentation map and key point coordinates. Next, the key point changes are calculated to obtain a first sequence. The third image and the first sequence are weighted and fused, and analyzed by a multi-head LSTM to obtain a first vector, a second vector, and a posture label. Finally, based on the first vector, the second vector, and the posture label, the optimal background is selected through a multi-feature matching mechanism.

[0005] To achieve the above object, the present invention provides the following technical solutions: An optimization method for an image recognition and motion tracking system of a touch interaction terminal for cultural relics and museums, comprising: Obtaining a user multi-modal image, where the multi-modal image includes a color image, a depth image, and an infrared image; extracting features from the multi-modal image to obtain a first image; Evaluating the quality of the multi-modal image to generate weights, weighting the first image to obtain a second image, fusing and decoding the second image features by using a feature pyramid and a decoder to obtain a third image, and passing the third image through a segmentation head and a key point detection head to obtain a segmentation image and key point coordinates; Calculate the change in the key point coordinates within the time window of the most recent N frames to obtain the first sequence; Weightedly fuse the third image and the first sequence to obtain the second sequence, and analyze the second sequence through LSTM to obtain the first vector, the second vector, and the pose label; Use the multi-feature matching mechanism to evaluate the first vector, the second vector, and the pose label, and select the optimal background.

[0006] Preferably, the process of obtaining the user's multi-modal image includes: Obtain the ambient light intensity using a light sensor, obtain a depth image using a depth camera, calculate the centroid position of the user, obtain the centroid depth distance, define a variable-size rectangular key area using the mapped centroid and distance, analyze the actual brightness information of the key area in the color image and compare it with the system-predefined target brightness level, and adjust the exposure time and gain using the intelligent light sensing system inside the camera according to the comparison result to obtain the multi-modal image.

[0007] Preferably, the process of evaluating the quality of the multi-modal image to generate weights and weighting the first image to obtain the second image includes: dividing the multi-modal image into image blocks, using local quality evaluation to real-time evaluate the signal fidelity index of the multi-modal image blocks, and generating respective local quality scores; calculating local fusion weights based on the local quality scores, and using the local fusion weights to perform a weighting process on the first image to obtain the second image.

[0008] Preferably, the process of performing feature fusion and decoding on the second image using the feature pyramid and the decoder includes: inputting the second image into the feature pyramid for fusion, where the feature pyramid is a hierarchical feature pyramid structure, and the feature pyramid that fuses the shallow second image processes the second image using respective independent weights; the feature pyramid that fuses the deep second image processes the second image using cross-modal shared weights; inputting the fused image into the Transformer decoder for decoding to obtain the third image, where the independent weights and the shared weights are learned during the model training process.

[0009] Preferably, the process of calculating the change in the key point coordinates within the time window of the most recent N frames to obtain the first sequence includes: corresponding the key point coordinates of the most recent N frames with the timestamps of the corresponding frames, and calculating the coordinate changes between adjacent frames.

[0010] Preferably, the process of weightedly fusing the third image and the first sequence to obtain the second sequence includes: calculating dynamic weights in real-time based on the third image, the first sequence, and relative position encoding; using the dynamic weights to perform a weighting process on the first sequence and the third image to generate the second sequence.

[0011] Preferably, the analysis of the second sequence by the LSTM includes: the LSTM uses a multi-head LSTM, and based on the original LSTM, an attention mechanism, multiple classification heads, a prediction head, and a current state representation head are added to process the LSTM hidden layer features respectively to obtain a first vector, a second vector, and a pose label.

[0012] Preferably, the evaluation process of the multi-feature matching mechanism includes: performing multi-dimensional weighting on the first vector according to the pose label to obtain a user vector, calculating the vector similarity between the user vector and the candidate background preset content vector to obtain a first matching result; based on the pose label, calculating the semantic similarity with the candidate background label by using semantic embedding technology to obtain a second matching result; combining the first matching result and the second matching result to determine the optimal matching background.

[0013] Preferably, the combination of the first matching result and the second matching result to determine the optimal matching background includes: setting an interaction threshold, calculating the action consistency coefficient in real time, when the action consistency coefficient is less than the interaction threshold, it is determined that the current continuous action has low consistency, and the weight of the first matching result is reduced, the weight of the second matching result is increased, and the background with the highest score is selected as the optimal matching background; if the action consistency coefficient is higher than the interaction threshold, calculate the matching degree between the candidate background and the second vector to obtain a forward-looking matching score; use the forward-looking matching score as a correction factor to adjust the weights of the first matching result and the second matching result, and select the background with the highest score as the optimal matching background.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By adopting color, depth, and infrared multi-modal inputs and combining local quality evaluation to perform weighted fusion on multi-modal images, it effectively suppresses the interference of low-quality data, can better cope with challenges such as illumination changes and occlusions, significantly improves the accuracy and robustness of user segmentation and feature extraction, avoids a series of subsequent interaction incoordination problems caused by recognition errors, and improves the stability of the user interaction experience from the source.

[0015] 2. By effectively combining spatial features and temporal features through weighted fusion and using multi-head LSTM for temporal modeling to generate user vectors and pose labels, it can better understand the continuous actions and internal states of users, its interaction response can better keep up with the rhythm and intentions of users, provides more natural, smooth, and user-expected interaction feedback, significantly improves the response sense and intelligence sense during the interaction process, and thus improves the user experience.

[0016] 3. A multi - feature matching mechanism that simultaneously considers continuous user vectors and discrete pose tags is adopted, and an action consistency coefficient is introduced for dynamic adjustment, ensuring that the background selection is not only morphologically relevant, but also highly consistent with the user's intention and context, effectively suppressing invalid switching in continuous interactions, significantly enhancing the continuity, stability, and fluency of visual presentation, and improving the interaction immersion. Brief Description of the Drawings

[0017] Figure 1 It is a flowchart of the method provided by an embodiment of the present invention; Figure 2 It is a flowchart of the hierarchical feature pyramid provided by an embodiment of the present invention; Figure 3 It is a flowchart of the multi - head LSTM processing provided by an embodiment of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Please refer to Figures 1 to 3 , the present invention provides an optimization method for the image recognition and action tracking system of a touch - interactive terminal for cultural relics and museums, and the technical solutions are as follows.

[0020] Embodiment 1

[0021] In order to improve the fluency of users' experience with the touch - interactive terminal in the museum and the users' experience, for this purpose, Museum A applies the optimization method for the image recognition and action tracking system of the touch - interactive terminal for cultural relics and museums. Figure 1 It is a flowchart of the method provided by an embodiment of the present invention.

[0022] As Figure 1 shown, the optimization method for the image recognition and action tracking system of the touch - interactive terminal for cultural relics and museums includes: Obtain multi - modal images of the user, where the multi - modal images include color images, depth images, and infrared images; perform feature extraction on the multi - modal images to obtain the first image; Evaluate the quality of the multi - modal images to generate weights, weight the first image to obtain the second image, use a feature pyramid and a decoder to perform feature fusion and decoding on the second image to obtain the third image, and obtain the segmentation image and key point coordinates by passing the third image through a segmentation head and a key point detection head; Calculate the change in key point coordinates within the time window of the most recent N frames to obtain the first sequence; Fuse the third image and the first sequence with weights to obtain a second sequence, and analyze the second sequence through an LSTM to obtain a first vector, a second vector, and a pose label; Evaluate the first vector, the second vector, and the pose label using a multi-feature matching mechanism to select the optimal background.

[0023] Specifically, a user multi-modal image is obtained in real time through a visible light camera, a depth camera, and an infrared camera in a high-definition camera. The multi-modal image includes a color image, a depth image, and an infrared image. The process of obtaining the user multi-modal image includes: First, obtain the ambient light intensity according to a light sensor, and at the same time obtain the user depth image according to the depth camera. Use the depth thresholding method to separate the user from the background, extract the user foreground, calculate the centroid position of the user foreground, obtain the depth value of the centroid position, and obtain the user distance d. Then, use the intelligent light sensing system inside the camera for intelligent adjustment. The core of the intelligent light sensing system is to analyze the real-time image captured by the visible light camera, and use the previously obtained user foreground area information and the user distance d to determine the key area for image analysis. The key area is to map the centroid into the color image, and with the centroid as the center, determine a rectangular frame whose side length is inversely proportional to the distance d. The area inside the rectangular frame is the key area. Analyze the actual brightness information of the key area in the real-time image and compare it with the target brightness level preset by the system. The target brightness level preset by the system is calibrated according to the visual effect of typical user images. Combine the obtained ambient light intensity and dynamically and intelligently adjust the exposure time and gain of the visible light camera in the high-definition camera.

[0024] Through this intelligent light sensing system that combines ambient light estimation, precise focusing and analysis of the target using depth information, and image brightness feedback based on the key area, it is ensured that the captured user images can obtain a continuously optimized and properly exposed lighting effect, effectively avoiding overexposure or underexposure problems. This not only directly improves the visual experience but also provides a strong guarantee for the stability and accuracy of subsequent recognition and tracking processing links that rely on high-quality image input.

[0025] Furthermore, set a noise threshold and a noise ratio threshold, which are calibrated through experimental statistics. Perform gray-scale processing on the infrared image and the depth image to obtain a gray-scale image, and calculate the local variance of the gray-scale image; calculate the proportion of the local variance values in the gray-scale image that are greater than the noise threshold; if the proportion is greater than the noise ratio threshold, use median filtering to perform noise suppression processing on the gray-scale image. The specific calculation formula for median filtering noise suppression processing is as follows: ; Where, respectively represent the horizontal and vertical coordinates of the pixel points in the image, represent the pixel point the value after median filtering, represent the defined neighborhood window, represent the median operation, represent the pixel points within the neighborhood window.

[0026] If the said ratio is less than the noise ratio threshold, Gaussian filtering is used to perform noise suppression processing on the grayscale image. The specific formula for Gaussian filtering noise suppression processing is as follows: ; wherein, respectively represent the horizontal and vertical coordinates of the pixel points in the image, represent the pixel point the value after Gaussian filtering, represent the Gaussian kernel radius (take 3 ) represent the Gaussian kernel at the offset the weight at, represent the pixel point the grayscale value of the neighboring pixels, represent the pixel point neighboring pixels.

[0027] Furthermore, if the multimodal image is a color image, it is directly input into a lightweight backbone network (in this invention, MobileNetV3 is used) for feature extraction to obtain the first image. If the multimodal image is a depth image or an infrared image, after converting the multimodal image into a single channel, a lightweight backbone network is used for feature extraction to obtain the first image.

[0028] Furthermore, evaluate the quality of the multimodal image to generate weights, and weight the first image to obtain the second image, including: First, convert the multimodal image into a grayscale image, divide the grayscale image into s*s overlapping multimodal image blocks, and use the Sobel operator to calculate the gradient magnitude and gradient direction of each pixel in the multimodal block. The specific calculation formula is as follows: ; wherein, Q represents the gradient energy score, and the larger the value, the clearer the image, respectively represent the horizontal and vertical coordinates of the pixel points in the image, N represents the side length pixels of each divided region block, represent the horizontal gradient within the region, represent the vertical gradient within the region.

[0029] Further, calculate the intrinsic quality index of the multimodal image. If the multimodal image is a depth image, calculate the proportion of valid pixels and normalize it. If the multimodal image is a color image, calculate the local RMS contrast and the proportion of non-saturated pixels and normalize them, and calculate the average value of the normalized local RMS contrast and the proportion of non-saturated pixels as the color image index. If the multimodal image is an infrared image, calculate the local standard deviation and normalize it. The normalization uses the Sigmoid function to normalize to between [0, 1]. By calculating these fast indexes, a basic image quality score is provided for each image patch.

[0030] Further, calculate the guided gradient consistency score. Preset a color threshold and an infrared threshold. The color threshold can be obtained through experimental statistical analysis of color image samples, and the infrared threshold is determined based on similar statistical analysis of infrared image samples, which are respectively used to judge whether there are significant visual structures or heat gradients. If the multimodal image is a color image, compare the gradient magnitude of the multimodal image patch with the color threshold. If it is greater than the color threshold, judge whether there is a significant visual change at the pixel position. When the gradient magnitude is greater than the color threshold, it is considered that there is a visually visible structure here. If the multimodal image is an infrared image, compare the gradient magnitude of the multimodal image patch with the infrared gradient threshold. When the gradient magnitude is greater than the infrared threshold, it is considered that there is a significant heat gradient here. When the gradient magnitude of the multimodal image patch is greater than the infrared threshold and greater than the color threshold, judge whether the depth measurement value of the corresponding position multimodal image patch is valid. If the measurement value is valid, mark this area as a key area.

[0031] Further, select the pixels in the key area to calculate the similarity of different modal images. First, calculate the cosine similarity using the gradient directions of the calculated multimodal images respectively, and then calculate the consistency score of each modality at this pixel. The consistency score of each modality is obtained by averaging its direction similarity with the other two modalities. Multiply the consistency score by the basic image quality score to assign a lower value to the non-key area. Then, obtain the final score of each modal image patch. Next, based on the final local quality scores of each modality, through cross-modal normalization using the Softmax function, calculate the dynamic fusion weights corresponding to the multimodal image patch. These weights will then be used to weight the first image to obtain the second image, where the average weight is calculated for the overlapping part.

[0032] Table 1 Dynamic Weighting Table for the Same Image Patch of Each Modality

[0033] Table 1 shows the dynamic weights given for the same image patch (randomly selecting the same image patch) in different environments. It can be seen from Table 1 that the weighted changes of local quality assessment for different modalities in different environments significantly improve the feature extraction ability for complex scenes and scenes with severe environmental interference.

[0034] Local quality assessment generates refined local reliability scores for each modality data through the intrinsic signal fidelity, inter-modal structural consistency, and physical rationality of the local regions of real-time multi-modal images. These scores are then used to dynamically calculate local fusion weights, enabling the subsequent feature fusion process to adaptively focus on using the modality information with higher scores and suppress the data sources with lower scores, thus significantly improving the overall quality and robustness of multi-modal fusion features in complex and noisy environments.

[0035] Furthermore, feature fusion is performed on the second image through a feature pyramid, and the fused image is decoded using a Transformer decoder to obtain a third image. The flowchart of the feature pyramid is as Figure 2 shown. The feature pyramid is a hierarchical feature pyramid structure, including an independent weight layer and a shared weight layer. The feature pyramid layers that fuse the shallow second image process the second image using their respective independent weights; the feature pyramid layers that fuse the deep second image process the second image using cross-modal shared weights; where the independent weights and the shared weights are learned during the model training process. The fused second image is input into the Transformer decoder for decoding to obtain a third image.

[0036] Furthermore, the model training process includes: Collect user operation images and label the user regions and key point regions (such as finger key points, mouth key points) in the user operation images to obtain a user operation data set; Input the user operation data set into an adaptive weighted multi-output segmentation model, and the segmentation model outputs the predicted values for the user operation data set; Calculate the loss value using the predicted values, the user regions, and the key point regions with a loss function; Calculate the gradient information using the backpropagation method for the loss value; update the network parameters of the segmentation model using the Adam optimizer according to the gradient information; If the number of training rounds of the segmentation model is equal to the preset number of rounds, then end the training of the segmentation model, otherwise repeat the above steps.

[0037] Furthermore, the third image is processed through a segmentation head and a key point detection head to obtain a segmented image and the coordinates of multiple user-defined key points.

[0038] Feature fusion of the second image feature through a hierarchical feature pyramid can not only better fuse multi-modal feature maps, but also significantly reduce the computational complexity of the feature pyramid layers, enhance the real-time performance of the computation, and use a Transformer to decode the second image to obtain a third image, making the image features clearer.

[0039] Further, a first sequence is generated by calculating the key point information within the time window of the most recent N frames. This sequence will contain the coordinates of the key points and their changes. First, a time window containing the most recent N frames is defined (the choice of N affects the time scale of the analysis and needs to be determined according to the application scenario). For each frame within the window, its precise timestamp and all key point coordinates need to be obtained, and ensure that the two strictly correspond. Then, all adjacent frames within the window are traversed (a total of N - 1 pairs, such as frame i and frame i + 1). For each pair of adjacent frames, the position change of the same key point (if it exists in both frames) is calculated to obtain a displacement vector, which reflects the movement direction and amplitude of the key point within this time interval. Finally, the data of these N - 1 frame intervals are organized in chronological order to form the first sequence. Among them, each interval (from frame i to frame i + 1) corresponds to a data unit that encapsulates three parts of information: (1) a set of displacement vectors of all key points calculated within this interval; (2) a set of key point coordinates of the starting frame (frame i); (3) a set of key point coordinates of the ending frame (frame i + 1).

[0040] By calculating the changes in the key point coordinates within the time window of the most recent N frames, dynamic information can be better captured, and temporal context features can be provided for the dynamic information, enhancing the representation of the user's motion features and facilitating improving the accuracy of subsequent tasks.

[0041] Further, the third image and the first sequence are fused with weights to obtain a second sequence, including: First, the key point coordinates corresponding to the time steps in the first sequence are mapped to a polar coordinate system with the center of the segmentation region associated with the third image feature as the origin to obtain a first relative position encoding with angle information and distance information, which is used to characterize the spatial distribution of the key points relative to the center of a specific region. Then, the features of the third image are extracted to obtain a first spatial vector, the displacement vectors of the key points stored in the first sequence are extracted to obtain a first key point vector, and an independent linear layer is used to map the first spatial vector and the first key point vector to a unified hidden dimension to obtain a second spatial vector and a second key point vector. Finally, the first relative position encoding is processed through a linear layer to obtain a second relative position encoding.

[0042] Further, the second spatial vector, the second key-point vector, and the second relative position encoding are concatenated along the feature dimension to obtain a gated input vector. The gated input vector is input into a lightweight gated network composed of a linear layer and a Sigmoid activation function for calculation to obtain a dynamic weighting vector. The calculated dynamic weighting vector is used as an element-wise weight to perform weighted fusion with the second key-point vector and the second spatial vector to obtain a fused feature vector. The specific calculation formula is as follows: ; where represents the fused feature vector, represents the dynamic weighting vector, represents the second key-point feature vector, represents the second spatial feature vector.

[0043] Further, the above operation steps are repeatedly executed for each frame within the corresponding time window to obtain a series of fused feature vectors. These fused feature vectors are sorted in chronological order to obtain a second sequence.

[0044] By weighted-fusing the features in the third image and the first sequence, the temporal and spatial information is effectively combined, enabling better discrimination of the current action state. Additionally, the relative position encoding is added, enhancing the ability to understand corresponding complex scenarios.

[0045] Further, the second sequence is analyzed through LSTM to obtain a first vector, a second vector, and a pose label. The processing flow chart of the LSTM is as shown in Figure 3As shown. The LSTM uses a multi-head LSTM, adding an attention mechanism, multiple classification heads, a first vector output head, and a second vector output head based on the original LSTM. First, input the second sequence into the LSTM network layer. The LSTM network, relying on its gating mechanism, processes the input second sequence step by step in time, effectively capturing short-term and long-term temporal dependencies, dynamic patterns, and state evolutions in the sequence data. Then, obtain the hidden state sequence output by the LSTM layer and apply the attention mechanism to these hidden states. This mechanism calculates the attention scores for each time step and performs a weighted sum of the hidden state sequence according to the attention scores, where the attention mechanism is an additive attention based on a learnable query vector. Specifically, it calculates the correlation scores between a learnable query vector Q and each hidden state through a small feed-forward network, then uses the Softmax function to convert these scores into attention weights, and finally calculates the weighted sum of the hidden state sequence to generate a context vector that can represent the key information of the entire sequence. Next, input the context vector into the first vector output head, the second vector output head, and multiple classification output heads (composed of Softmax and fully connected layers) for processing to obtain a first vector, a second vector, and pose labels. The first vector represents the continuous and dynamic features of the user's actions over time, is a refined feature representation of the current moment state, and condenses the understanding of the complete user interaction information; the second vector represents the temporal dynamic law learned by the model from the second sequence and makes a predictive representation of the actions in the next frame; the pose label is a multi-dimensional label formed by predicting the categories and their probabilities of preset keywords in each dimension and combining the categories with the highest probabilities. (For example, [emotion: positive, gesture: thumbs up approval, intention: taking a photo pose, atmosphere: informal]).

[0046] By designing a multi-head LSTM and generating the required vectors and pose labels through multiple independent output heads, the user state can be characterized more comprehensively. At the same time, the lightweight fusion mechanism and key designs such as efficient LSTM units it adopts enable the entire tracking to have excellent real-time operation capabilities while maintaining powerful functions.

[0047] Further, a multi-feature matching mechanism is used to evaluate the first vector, the second vector, and the pose label to match the optimal background, including: First, obtain the first vector and the pose label from the output of the multi-head LSTM, and use the multi-feature matching mechanism to evaluate the first vector and the pose label. The steps of the multi-feature matching mechanism include analyzing various aspects of information contained in the pose label, and based on this information, evaluating their correlation with specific dimensions in the first vector. Based on the above correlation evaluation, dynamically adjust the weights of each dimension in the first vector to obtain a user vector. Among them, the dynamic adjustment of the weights of each dimension in the first vector is achieved by dynamically adjusting the weights using a predefined lookup table template. The predefined lookup table template is a predefined mapping relationship that establishes the association between the semantic categories (such as emotion categories, gesture categories, intention categories, etc.) contained in the pose label and the dimension groups in the first vector, as well as the corresponding weight adjustment strategies (such as increasing, decreasing, or maintaining weights). For example, if the user interaction label indicates that the user is in a positive mood and makes a like gesture, the system will increase the weights of the dimensions related to "facial smile degree" and "hand closing / specific shape", and at the same time reduce the weights of the dimensions related to irrelevant or contradictory features such as "palm opening degree". Then, use the cosine similarity to calculate the similarity between the user vector and the content vector preset for the background to obtain the first matching result, where the content vector preset for the background is the best content vector selected by statistically analyzing the background.

[0048] Further, convert the pose label into a natural language description with richer information to obtain a key information statement, and then apply semantic embedding technology to calculate the similarity. The semantic embedding technology is implemented using an advanced pre-trained language model in the field. Considering the real-time requirements, a lightweight or distilled Transformer model is preferred. This method uses a model based on the Sentence-BERT framework. First, map the key information statement and the semantic label preset for the candidate background to the same high-dimensional semantic space, and then calculate the semantic similarity between the key information statement and the semantic label of the candidate background in the semantic space to obtain the second matching result. Among them, the semantic label of the candidate background (for example, keywords describing the background theme or style) can be pre-annotated manually. Combine the first matching result and the second matching result to determine the optimal matching background.

[0049] By simultaneously considering the first vector and the pose label, the multi-feature matching mechanism can understand the user's current state more comprehensively and deeply. This makes the recommended background not only visually related to the action form, but also contextually and atmospherically in line with the inner meaning of the action, and the matching effect is more intelligent and appropriate.

[0050] Further, when determining the optimal matching background by combining the first matching result and the second matching result, an action consistency coefficient is calculated in real time to quantify the stability and autocorrelation of the continuous action feature vectors of the user within the most recent N-frame time window. The specific calculation formula is as follows: ; where represents the action consistency coefficient at the current time t, represents the sensitivity parameter, which is used to adjust the sensitivity of the action consistency, represents the size of the time window, represents at the th frame of the action feature vector, represents at the th frame of the action feature vector, represents the exponential function with base e.

[0051] Further, an interaction threshold is set. The interaction threshold is based on the statistical analysis of the action consistency coefficient under different user interaction modes (such as, stably maintaining a posture, posture conversion, aimless movement, etc.), and is optimized and selected in combination with the feedback results of the user experience test, aiming to achieve the best interaction fluency perceived by the user. When the action consistency coefficient is less than the interaction threshold, it is determined that the current continuous action has low consistency, and the first vector sequence will become unstable, resulting in a reduction in the reliability of the first matching evaluation result. Therefore, at this time, the weight of the first matching result will be reduced, and the weight of the second matching result will be increased. The lower the action consistency coefficient, the smaller the weight of the first matching result, and the greater the weight of the second matching effect. The way of weight reduction has a linear relationship with the action consistency coefficient, and the linear ratio is adjusted through the sensitivity coefficient . The larger the value, the greater the reduction amplitude of the weight of the first matching result for each unit decrease in the consistency coefficient . Through this adjustment, more attention can be paid to the user's current action, preventing the interference of invalid actions.

[0052] Further, when the action consistency coefficient is greater than the interaction threshold, the average of the first matching result and the second matching result calculated for the current frame is taken to obtain a basic comprehensive score reflecting the current state matching degree. Next, the cosine similarity between the second vector and the preset vector of the background is calculated as the prospective matching score. Then, the prospective matching score is used as a correction factor to adjust the basic comprehensive score to obtain the final comprehensive score for ranking. The specific adjustment formula is as follows: ; where, represents the final comprehensive score, represents the basic comprehensive score, represents the forward-looking matching score, is a preset positive adjustment intensity coefficient, which controls the overall correction range of the forward-looking matching score to the basic score. The value of this coefficient is preset according to application requirements.

[0053] By calculating the action consistency coefficient, it can effectively prevent the matching of irrelevant actions of the user in continuous matching. In addition, using the action consistency coefficient and the prediction-based correction process can more comprehensively reflect the intention and trend of the user's continuous actions, improve the accuracy of continuous background matching, and greatly optimize the user's interaction experience during shooting.

[0054] In addition, the method of the present invention also supports the user to manually select the background for shooting through the touch interaction terminal according to their own needs. The method for optimizing the image recognition and action tracking system of the touch interaction terminal for cultural relics proposed by the present invention effectively overcomes the interference of environmental factors on feature extraction and enhances the understanding and tracking of the user's continuous actions and internal intentions. Compared with the traditional background matching method, the present invention significantly improves the accuracy, robustness and fluency of automatic background matching of the touch interaction terminal for cultural relics in complex scenarios. In addition, the present invention also supports the user to manually select the background, taking into account both intelligent recommendation and the user's right to independent choice, and providing a better experience and interactivity.

[0055] Embodiment 2

[0056] As the user's enthusiasm for culture and history increases, the user's demand for museum interaction devices becomes higher. Museum B has relatively high requirements for the recognition accuracy of user postures and action details, and the computing resource configuration of the touch interaction terminal is relatively abundant. Therefore, Museum B introduces the method for optimizing the image recognition and action tracking system of the touch interaction terminal for cultural relics provided by the present invention in the touch interaction device terminal. The specific implementation method is as follows: First, the exposure time and gain inside the camera are adjusted through the intelligent light sensing system inside the camera to obtain multi-modal images, and then noise suppression operations are performed on the multi-modal images. Feature extraction is performed on the denoised images. Since Museum B has relatively high requirements for image recognition and action tracking, and the computing resource configuration of the touch interaction terminal is relatively abundant, the backbone network for feature extraction in this embodiment selects the EfficientNet-B1 backbone network. If the multi-modal image is a color image, it is directly input into the backbone network for feature extraction to obtain the first image. If the multi-modal image is a depth image or an infrared image, the multi-modal image is converted into a single channel and then the backbone network is used for feature extraction to obtain the first image.

[0057] Further, the multi-modal image is divided into image patches, and the signal fidelity index of the multi-modal image patches is evaluated in real time using local quality assessment to generate respective local quality scores; local fusion weights are calculated based on the local quality scores, and the first image is weighted using the local fusion weights to obtain a second image. Then, the second image is input into a feature pyramid for fusion. In this embodiment, the feature pyramid uses a bidirectional information flow path pyramid BiFPN with better feature fusion effect. BiFPN receives the quality-weighted feature maps from different stages of the backbone network in the second image as inputs. Through its internal multiple bidirectional fusion levels, BiFPN can efficiently fuse feature information of different resolutions. During the fusion process, a weighted feature fusion mechanism is adopted, where the contribution degree of different input features to the fusion result is determined by learnable weight parameters to achieve a more adaptive feature combination. After being processed by the BiFPN, a set of fused feature maps with rich multi-scale information is output. The feature maps are decoded through a Transformer decoder to obtain a third image, and the third image is input into a segmentation head and a key point detection head to obtain a segmented image and key point coordinates. Further, the change in the key point coordinates within the most recent N-frame time window is calculated to obtain a first sequence. The first sequence, the third image, and the relative position encoding are weighted and fused through a lightweight gated network to obtain a second sequence. Then, the second sequence is input into a multi-head LSTM for analysis to obtain a first vector, a second vector, and a pose label. The first vector is weighted multi-dimensionally according to the pose label to obtain a user vector. The vector similarity between the user vector and the candidate background preset content vector is calculated to obtain a first matching result; based on the pose label, the semantic similarity with the candidate background label is calculated using semantic embedding technology to obtain a second matching result; then, the action consistency coefficient is calculated in real time. When the action consistency coefficient is less than the interaction threshold, it is determined that the current continuous action has low consistency, the weight of the first matching result is reduced, the weight of the second matching result is increased, and the background with the highest score is selected as the optimal matching background; if the action consistency coefficient is higher than the interaction threshold, the matching degree between the candidate background and the second vector is calculated to obtain a forward-looking matching score; the forward-looking matching score is used as a correction factor to adjust the weights of the first matching result and the second matching result, and the background with the highest score is selected as the optimal matching background.

[0058] To verify the advantages of introducing the method of the present invention into Museum B, a set of comparative experiments were conducted. The experimental group was the image recognition and action tracking optimization method for the cultural and museum touch interaction terminal proposed by the present invention, and the control group was the existing method of Museum B. The collected test data set included a variety of representative user scenarios, covering test data sets of different user actions, different environmental illuminations, and different scenario types. The best-matched background was marked for the test data set, and the background matching accuracy rate, false matching probability, and missed matching probability were calculated and compared. The background matching accuracy rate represents the probability that the background in the test set is matched to the marked background, and the false matching probability is the probability that the background match in the test set does not match the marked background. In addition, the user satisfaction score was added. The user satisfaction was obtained through the subjective evaluation of the user's interaction experience. 100 users who had used both methods were selected for a questionnaire survey.

[0059] It can be clearly seen from Table 2 that the method proposed by the present invention has significantly improved the matching accuracy rate compared with the existing method in Museum B, reduced the false matching probability, and also greatly improved in terms of user satisfaction, indicating that the method of the present invention has significantly improved the user's interaction experience.

[0060] Table 2 Data table of comparative experiment results

[0061] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. Optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal, characterized in that, Including: Obtain a user's multimodal image, where the multimodal image includes a color image, a depth image, and an infrared image; perform feature extraction on the multimodal image to obtain a first image; Evaluate the quality of the multimodal image to generate weights, weight the first image to obtain a second image, use a feature pyramid and a decoder to perform feature fusion and decoding on the second image to obtain a third image, and obtain a segmentation image and key point coordinates by passing the third image through a segmentation head and a key point detection head; Calculate the change in key point coordinates within the most recent N-frame time window to obtain a first sequence; Weightedly fuse the third image and the first sequence to obtain a second sequence, and analyze the second sequence through an LSTM to obtain a first vector, a second vector, and a pose label; Use a multi-feature matching mechanism to evaluate the first vector, the second vector, and the pose label, and select the optimal background.

2. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 1, characterized in that, The process of obtaining the user's multimodal image includes: Obtain the ambient light intensity using a light sensor, obtain a depth image using a depth camera, calculate the centroid position of the user, obtain the centroid depth distance, use the mapped centroid and distance to define a variable-size rectangular key area, analyze the actual brightness information of the key area in the color image and compare it with the system-predefined target brightness level, and according to the comparison result, use the intelligent light sensing system inside the camera to adjust the exposure time and gain to obtain a multimodal image.

3. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 2, characterized in that, The evaluating the quality of the multimodal image to generate weights and weighting the first image to obtain a second image includes: dividing the multimodal image into image blocks, using local quality assessment to real-time evaluate the signal fidelity index of the multimodal image blocks to generate respective local quality scores; calculating local fusion weights based on the local quality scores, and using the local fusion weights to perform a weighting process on the first image to obtain a second image.

4. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 3, characterized in that, The using a feature pyramid and a decoder to perform feature fusion and decoding on the second image includes: inputting the second image into the feature pyramid for fusion, where the feature pyramid is a hierarchical feature pyramid structure, and the feature pyramid that fuses the shallow second image processes the second image using respective independent weights; the feature pyramid that fuses the deep second image processes the second image using cross-modal shared weights; inputting the fused image into a Transformer decoder for decoding to obtain a third image, where the independent weights and the shared weights are learned during the model training process.

5. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 1, characterized in that, The calculating the change in key point coordinates within the most recent N-frame time window to obtain a first sequence includes: corresponding the key point coordinates of the most recent N frames with the timestamps of the corresponding frames, and calculating the coordinate changes between adjacent frames.

6. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 1, characterized in that, The weightedly fusing the third image and the first sequence to obtain a second sequence includes: extracting the spatial features corresponding to the third image and the key point features corresponding to the first sequence, concatenating the spatial features, the key point features, and the relative position encoding along the feature dimension, and using a lightweight gating network to real-time calculate dynamic weights; using the dynamic weights to perform a weighting process on the spatial features and the key point features to generate a second sequence.

7. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 6, characterized in that, The analysis of the second sequence by the LSTM includes: the LSTM uses a multi-head LSTM, and based on the original LSTM, an attention mechanism, multiple classification heads, a first vector output head, and a second vector output head are added to process the LSTM hidden layer features respectively to obtain a first vector, a second vector, and a pose label.

8. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 7, characterized in that, The evaluation process of the multi-feature matching mechanism includes: performing multi-dimensional weighting on the first vector according to the pose label to obtain a user vector, calculating the vector similarity between the user vector and the candidate background preset content vector to obtain a first matching result; based on the pose label, using semantic embedding technology to calculate the semantic similarity with the candidate background label to obtain a second matching result; combining the first matching result and the second matching result to determine the optimal matching background.

9. The optimization method for image recognition and motion tracking system of cultural relics touch interactive terminal according to claim 8, characterized in that, The combination of the first matching result and the second matching result to determine the optimal matching background includes: setting an interaction threshold, calculating the action consistency coefficient in real time. When the action consistency coefficient is less than the interaction threshold, it is determined that the current continuous action has low consistency, and the weight of the first matching result is reduced, and the weight of the second matching result is increased, and the background with the highest score is selected as the optimal matching background; if the action consistency coefficient is higher than the interaction threshold, calculate the matching degree between the candidate background and the second vector to obtain a forward-looking matching score; use the forward-looking matching score as a correction factor to adjust the weights of the first matching result and the second matching result, and select the background with the highest score as the optimal matching background.