AI interview English body language aided teaching method and system fused with computer vision

By introducing dynamic color gamut difference threshold and multi-frame timing prediction in human body segmentation technology, the problem of low human body segmentation accuracy under virtual background is solved, and higher segmentation accuracy and more accurate posture estimation are achieved.

CN120147933AInactive Publication Date: 2025-06-13JINING POLYTECHNIC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510289142.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has low accuracy in human segmentation under virtual backgrounds, especially in key parts such as arms and background separation is not accurate enough, which affects the accuracy of overall posture estimation and the evaluation effect of non-verbal communication.

Method used

By dynamically adjusting the color gamut difference threshold and multi-frame timing prediction, the human body segmentation accuracy is improved. The specific steps include: obtaining the interview video stream under the virtual background, extracting the RGB image and background mask image, inputting the segmentation network to generate the initial human body segmentation results, predicting the current frame limb movement trend based on the previous multi-frame results, strengthening the edge contour with the weight of the heat map of the human body key points, and eliminating abnormal fragments to obtain the correct segmentation results.

Benefits of technology

It significantly improves the accuracy of human segmentation under virtual background interference, reduces confusion of limb edge segmentation, improves the accuracy of overall posture estimation and the evaluation effect of non-verbal communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147933A_ABST
    Figure CN120147933A_ABST
Patent Text Reader

Abstract

The invention discloses an AI interview English limb language aided teaching method and system fused with computer vision, particularly relates to the technical field of image processing, and is used for solving the problems of low limb segmentation precision and disjunction of action evaluation and culture scenes caused by existing virtual background interference. Obtaining an interview video stream under a virtual background through a camera, and extracting an RGB image and a background mask image; generating a high-precision human body segmentation result based on dynamic color gamut threshold adjustment and multi-frame motion prediction; combining the weight of the thermodynamic diagram to strengthen the head-hand contour and remove abnormal fragments; aligning the skeleton track with the voice timestamp, and extracting a semantic association action fragment; and after screening culture sensitive actions, matching a cross-culture rule base to generate an adaptive suggestion. According to the method, through background type adaptive segmentation, cross-modal space-time alignment and culture rule dynamic adaptation, the segmentation accuracy and evaluation precision are remarkably improved, and efficient and culture-compatible body language optimization feedback is provided for remote interview.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more specifically, to an AI interview English body language assisted teaching method and system integrating computer vision. Background Art

[0002] In remote interviews and online education scenarios, computer vision technology is widely used to capture and analyze users' body language, especially in image processing and human posture estimation based on deep learning. In the AI ​​interview English body language assisted teaching system, user images are acquired through the camera, and then image segmentation, feature extraction and pattern recognition technologies are used to accurately capture and analyze the interviewee's gestures, movements and expressions, thereby providing a basis for the evaluation and feedback of non-verbal communication.

[0003] In the existing technology, since users often use virtual backgrounds (such as Zoom green screen) in remote interviews, there is confusion in the segmentation of limb edges, especially in the separation of key parts (such as arms) from the background, which is not accurate enough, thus affecting the accuracy of overall posture estimation, and further reducing the evaluation effect of the AI ​​interview English body language assisted teaching system on non-verbal communication. Summary of the invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an AI interview English body language assisted teaching method and system integrating computer vision to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] The AI ​​interview English body language assisted teaching method integrated with computer vision includes the following steps:

[0007] S1. Obtain the interview video stream under the virtual background through the camera, and extract the RGB image and background mask image of the video frame;

[0008] S2, input the RGB image and the background mask image into the segmentation network to generate the initial human segmentation result of the current frame;

[0009] S3, predicting the limb movement trend of the current frame based on the initial human body segmentation results of the previous multiple frames, comparing the displacement deviation between the predicted contour and the actual segmentation edge, and marking the abnormal fragments with excessive dynamic thresholds;

[0010] S4, combining the weights of the key points of the human body heat map to strengthen the edge contours of the head and hands, and remove abnormal fragments to obtain the corrected segmentation results;

[0011] S5, aligning the human skeleton trajectory of the corrected segmentation result with the synchronized English speech timestamp, and extracting speech semantic keywords;

[0012] S6. Screen the culturally sensitive action segments according to the relevance between the semantic keywords and the skeletal trajectories;

[0013] S7. Match the culturally sensitive action segments with the gesture amplitude and sitting posture standards in the cross-cultural rule library to generate body language suggestions adapted to the enterprise scenario.

[0014] In a preferred embodiment, an interview video stream under a virtual background is obtained through a camera, and the RGB image and the background mask image of the video frame are extracted, including:

[0015] S1a. Real-time collect the interview video stream containing the virtual background through the camera;

[0016] S1b. Decode the interview video stream frame by frame and extract the RGB image of each video frame;

[0017] S1c. Perform chroma key matte processing on the RGB image to generate a binary background mask image corresponding to the virtual background area;

[0018] S1d. Store the RGB image and the corresponding background mask image into the video frame buffer queue.

[0019] In a preferred embodiment, the RGB image and the background mask image are input into a segmentation network to generate the initial human body segmentation result of the current frame, including:

[0020] S2a. Input the RGB image and the background mask image into the segmentation network, and extract the color distribution features of the background mask image. The color distribution features include the color level mean and the dispersion degree of the main color channel;

[0021] S2b. Classify the virtual background type as a static single-color background or a dynamic multi-color background based on the color level mean and the dispersion degree of the main color channel;

[0022] S2c. Dynamically adjust the color gamut difference threshold according to the virtual background type: if it is a static single-color background, the color gamut difference threshold is set to a fixed range; if it is a dynamic multi-color background, the color gamut difference threshold linearly expands as the dispersion degree of the main color channel increases;

[0023] S2d. Perform pixel-level filtering on the RGB image based on the adjusted color gamut difference threshold to generate the initial human body segmentation result of the current frame.

[0024] In a preferred embodiment, based on the initial human body segmentation results of the previous multiple frames, predict the limb movement trend of the current frame, compare the displacement deviation between the predicted contour and the actual segmentation edge, and mark those exceeding the dynamic threshold as abnormal fragments, including:

[0025] S3a. Obtain the initial human body segmentation results of the previous multiple frames, extract the displacement vectors of human body key points in each frame, and generate the key point motion trajectories;

[0026] S3b. Predict the limb motion trends of the current frame based on the key point motion trajectories, and generate the expected limb contours;

[0027] S3c. Calculate the pixel-level displacement deviation between the actual segmentation edge of the current frame and the expected limb contour. If the pixel-level displacement deviation exceeds the dynamic displacement threshold, mark the corresponding area as abnormal debris;

[0028] S3d. The dynamic displacement threshold is dynamically adjusted according to the average displacement deviation of the previous multiple frames.

[0029] In a preferred embodiment, the head and hand edge contours are strengthened by combining the weights of the human body key point heat maps, and the abnormal debris is removed to obtain the corrected segmentation result, including:

[0030] S4a. Generate multi-scale Gaussian heat maps based on the human body bone key point coordinates. Use small-scale high-weight Gaussian kernels for the head and hand key points, and use large-scale low-weight Gaussian kernels for the torso key points;

[0031] S4b. Perform region adaptive fusion of the multi-scale Gaussian heat maps and the initial human body segmentation results: Strengthen the edge gradients of the head and hand regions according to the heat map weights, and retain the original contours of the torso regions according to the initial segmentation results;

[0032] S4c. Extract the binary mask of the interference region according to the abnormal debris marking result, and use the heat map weight-guided local region growing algorithm to fill the holes in the interference region;

[0033] S4d. Perform morphological closing operations with edge constraints on the non-interference regions, and perform anisotropic diffusion filtering on the interference regions after region growing to generate the corrected segmentation result.

[0034] In a preferred embodiment, anisotropic diffusion filtering is performed on the interference regions after region growing, and the diffusion coefficient is adjusted according to the heat map weight value: ; where represents the anisotropic diffusion coefficient at the pixel coordinate , represents the gradient magnitude of the fused segmentation map at the coordinate , represents the gradient threshold parameter.

[0035] In a preferred embodiment, align the human body bone trajectory of the corrected segmentation result with the synchronous English speech timestamps, and extract the speech semantic keywords, including:

[0036] S5a. Perform time series sampling on the human skeleton key point trajectories of the corrected segmentation results to generate the time series data of the displacements of the skeleton key points;

[0037] S5b. Perform phoneme-level segmentation on the synchronized English speech data to generate the sequence of the peak timestamps of the speech waveforms;

[0038] S5c. Align the time series data of the displacements of the skeleton key points and the sequence of the peak timestamps of the speech waveforms through the dynamic time warping algorithm to establish the skeleton-speech time mapping table;

[0039] S5d. Extract the semantic keywords that appear synchronously with the skeleton actions in the speech content based on the skeleton-speech time mapping table, and the semantic keywords include the vocabulary related to the target enterprise culture.

[0040] In a preferred embodiment, screening the culture-sensitive action segments according to the relevance between the semantic keywords and the skeleton trajectories, including:

[0041] S6a. Calculate the time-space overlap degree between the semantic keywords and the human skeleton key point trajectories to generate the cross-modal correlation degree score. The time-space overlap degree is the product of the displacement amplitude of the corresponding human skeleton key points during the time period when the semantic keywords appear and the semantic intensity of the speech content;

[0042] S6b. Dynamically adjust the cross-modal correlation degree score threshold according to the culture type to which the target enterprise belongs, and screen the candidate action segments whose cross-modal correlation degree scores exceed the threshold;

[0043] S6c. Perform culture sensitivity verification on the candidate action segments, and filter out the action segments that conflict with the target enterprise culture based on the predefined gesture amplitude threshold range and sitting posture uprightness score gradient in the cross-cultural body language rule library;

[0044] S6d. Integrate the verified action segments to generate the set of culture-sensitive action segments.

[0045] In a preferred embodiment, match the gesture amplitude and sitting posture standards of the cross-cultural rule library for the culture-sensitive action segments to generate the body language suggestions adapted to the enterprise scenario, including:

[0046] S7a. Obtain the action type labels and corresponding timestamps in the set of culture-sensitive action segments;

[0047] S7b. Query the gesture amplitude threshold range and sitting posture uprightness score gradient in the cross-cultural body language rule library according to the action type labels;

[0048] S7c. Compare and analyze the actual gesture amplitude in the culture-sensitive action segments with the gesture amplitude threshold range. If the actual gesture amplitude exceeds the threshold range, mark it as the action to be corrected;

[0049] S7d. Generate body language suggestions adapted to the target enterprise scenario based on the marked results of the actions to be corrected and the sitting posture rectitude scoring gradient. The suggestions include the adjustment range of gesture amplitude and the optimized sitting posture angle.

[0050] On the other hand, the present invention provides an AI interview English body language assisted teaching system integrating computer vision, including a virtual video acquisition module, an image segmentation generation module, a motion trend prediction module, a contour correction processing module, a voice-skeleton alignment module, a cultural action screening module, and a body language suggestion generation module;

[0051] Virtual video acquisition module: Obtain the interview video stream under the virtual background through a camera, and extract the RGB image and the background mask image of the video frame;

[0052] Image segmentation generation module: Input the RGB image and the background mask image into the segmentation network to generate the initial human body segmentation result of the current frame;

[0053] Motion trend prediction module: Predict the limb motion trend of the current frame based on the initial human body segmentation results of the previous multiple frames, compare the displacement deviation between the predicted contour and the actual segmentation edge, and mark those exceeding the dynamic threshold as abnormal fragments;

[0054] Contour correction processing module: Strengthen the head and hand edge contours by combining the weights of the human key point heat maps, and remove the abnormal fragments to obtain the corrected segmentation result;

[0055] Voice-skeleton alignment module: Align the human skeleton trajectory of the corrected segmentation result with the synchronous English speech timestamp, and extract the voice semantic keywords;

[0056] Cultural action screening module: Screen the culturally sensitive action segments according to the relevance between the semantic keywords and the skeleton trajectory;

[0057] Body language suggestion generation module: Match the culturally sensitive action segments with the gesture amplitude and sitting posture standards in the cross-cultural rule library to generate body language suggestions adapted to the enterprise scenario.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] 1. The present invention significantly improves the accuracy of human body segmentation under virtual background interference through a technical means that combines dynamic adjustment of the color gamut difference threshold with multi-frame temporal prediction. Traditional methods cannot adapt to the color gamut changes of dynamic backgrounds due to fixed thresholds, resulting in confusion in the segmentation of limb edges, especially a high mis-segmentation rate in key parts such as the arms. The present invention introduces a background type classification mechanism into the segmentation network, adaptively adjusts the color gamut difference threshold according to the static or dynamic characteristics of the virtual background (such as expanding the threshold range when the dispersion of the dynamic background increases), effectively suppressing the interference of background noise on the segmentation result; at the same time, predicts the limb movement trend of the current frame based on the segmentation results of the previous multiple frames, and marks abnormal fragments through displacement deviation detection, solving the problem of contour breakage caused by instantaneous noise in single-frame segmentation.

[0060] 2. The present invention realizes multi-dimensional optimization and cultural adaptability feedback of body language evaluation through cross-modal alignment and cross-cultural rule matching. Traditional methods only rely on the isolated analysis of body movements based on visual data, ignoring the spatio-temporal correlation between speech semantics and actions, resulting in the disconnection of evaluation results from the interview scenario. The present invention dynamically aligns the corrected human body bone trajectory with the English speech timestamp, extracts strongly correlated segments between semantic keywords and bone actions, and takes into account both time synchronization and spatial amplitude consistency when screening culturally sensitive actions; further combines the preset gesture amplitude and sitting posture standards in the cross-cultural rule library to generate body language suggestions adapted to the corporate culture. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a flowchart of the AI interview English body language assisted teaching method integrating computer vision according to the present invention;

[0062] Figure 2 is a schematic structural diagram of the AI interview English body language assisted teaching system integrating computer vision according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] Embodiment 1: Figure 1 An AI interview English body language assisted teaching method integrating computer vision according to the present invention is given, which includes the following steps:

[0065] S1. Obtain the interview video stream under the virtual background through a camera, and extract the RGB image and background mask image of the video frame.

[0066] S2. Input the RGB image and the background mask image into the segmentation network to generate the initial human segmentation result of the current frame.

[0067] S3. Predict the limb movement trend of the current frame based on the initial human segmentation results of the previous multiple frames, compare the displacement deviation between the predicted contour and the actual segmentation edge, and mark those exceeding the dynamic threshold as abnormal fragments.

[0068] S4. Strengthen the head and hand edge contours by combining the weights of the human key point heat maps, and remove the abnormal fragments to obtain the corrected segmentation result.

[0069] S5. Align the human bone trajectory of the corrected segmentation result with the synchronized English speech timestamp, and extract the speech semantic keywords.

[0070] S6. Screen the culture-sensitive action segments according to the relevance between the semantic keywords and the bone trajectory.

[0071] S7. Match the culture-sensitive action segments with the gesture amplitude and sitting posture standards in the cross-cultural rule library to generate limb language suggestions adapted to the enterprise scenario.

[0072] The camera captures the interview video stream with a virtual background in real time, decodes the interview video stream frame by frame, extracts the RGB image from each frame of the video stream. The pixel values of the RGB image contain the color gamut information of the red, green, and blue channels; perform chroma key matte processing on the RGB image. The chroma key matte processing is based on a preset background color threshold range, marks the pixels belonging to the virtual background area in the RGB image as the background, and generates a binary background mask image. The pixel value of the background area in the binary background mask image is 1, and the pixel value of the human body area is 0; store the RGB image and the corresponding binary background mask image in the video frame buffer queue. The video frame buffer queue stores continuous video frame data in chronological order for subsequent temporal analysis of multiple frames of data.

[0073] In the chroma key matte processing, if the virtual background is a green screen, the background color threshold range is the color level interval of the green channel; if the virtual background is a dynamic special effect background, the color threshold range is dynamically adjusted according to the real-time rendering characteristics of the dynamic background. For example, in the case of a dynamic particle special effect background, the color threshold range fluctuates periodically with the change of the particle color.

[0074] Among them, when generating the binary background mask image, perform a dilation operation on the pixels in the RGB image that match the background color threshold to fill the holes in the background area caused by uneven illumination and ensure the connectivity of the background mask image.

[0075] Input the RGB image and the background mask image into the segmentation network. The segmentation network receives the three-channel color gamut data of the RGB image and the binary region markers of the background mask image, and extracts the color distribution features of the background mask image. The color distribution features include the color level mean and the dispersion degree of the main color channel. The color level mean of the main color channel is the average color level value of the pixels marked as the background region in the background mask image in a certain channel (red, green, or blue) in the RGB color space, and the dispersion degree is the standard deviation of the color level values of this channel. Based on the color level mean and the dispersion degree of the main color channel, classify the virtual background type through the threshold determination logic: if the dispersion degree of the main color channel is less than or equal to a preset first threshold (for example, 5 color levels), then determine that the virtual background type is a static monochromatic background; if the dispersion degree is greater than the first threshold, then determine that the virtual background type is a dynamic multi-color background.

[0076] Dynamically adjust the color gamut difference threshold according to the classification result: when the virtual background type is a static monochromatic background, the color gamut difference threshold is set to a fixed range (for example, ±5 color levels), and the pixels whose deviation of the pixel color level from the color level mean of the main color channel is within this range in the RGB image are determined to be the background; when the virtual background type is a dynamic multi-color background, the color gamut difference threshold linearly expands with the increase of the dispersion degree of the main color channel. For example, when the dispersion degree is 10 color levels, the color gamut difference threshold expands to ±8 color levels, and when the dispersion degree is 15 color levels, it expands to ±12 color levels, so as to achieve the adaptive filtering of the dynamic background.

[0077] The adjusted color gamut difference threshold is used for pixel-level filtering of the RGB image. Calculate the difference between the color level of each pixel in the RGB image and the color level mean of the main color channel. If the difference value exceeds the current color gamut difference threshold, then determine that the pixel belongs to the human body region, otherwise mark it as the background region, and finally generate the initial human body segmentation result of the current frame.

[0078] In the dynamic multi-color background scenario, for example, when the user uses a dynamic virtual background with particle effects, the particle color changes gradually or flickers over time, the dispersion degree of the main color channel fluctuates periodically due to the particle color change, and the color gamut difference threshold is dynamically adjusted according to the real-time change of the dispersion degree to ensure the segmentation accuracy of the human body region under the interference of different color particles.

[0079] In the static monochromatic background scenario, for example, when the user uses a pure green curtain as the virtual background, the dispersion degree of the main color channel (green channel) is relatively low, and the color gamut difference threshold remains in a small range to avoid mis-segmentation of the edges of the arms or hair caused by color level fluctuations.

[0080] When generating the initial human body segmentation result, perform morphological closing operation on the pixels in the RGB image that match the background color gamut difference threshold to fill the small holes caused by background noise, and at the same time retain the edge details of the human body region.

[0081] When predicting the limb movement trend of the current frame based on the initial human segmentation results of the previous multiple frames, the displacement vectors of the human key points in each frame are extracted from the initial human segmentation results of the previous multiple frames. The human key points include the coordinate points of the head, shoulders, elbows, and wrists. The displacement vector represents the amount of position change of the same key point in the previous multiple frames and is calculated by the coordinate difference of the key points between adjacent frames. For example, if the coordinate of the elbow key point in the previous frame is (x1, y1) and the current frame is (x2, y2), then the displacement vector is (x2 - x1, y2 - y1).

[0082] Generate the key point movement trajectories based on the displacement vectors of the previous multiple frames. The key point movement trajectories are composed of the displacement vector sequences of the same key point in the previous multiple frames and are used to describe the movement direction and speed change law of the key point. For example, if the elbow key point moves uniformly in the upper right direction in five consecutive frames, the movement trajectory shows a linearly increasing trend.

[0083] When predicting the limb movement trend of the current frame based on the key point movement trajectories, the linear extrapolation method is used to extend the key point trajectories. The linear extrapolation method predicts the expected position of the key points in the current frame through the mean value and direction of the displacement vectors. For example, if the average displacement vector of the elbow key point in the previous five frames is (Δx, Δy), then the expected position in the current frame is (x5 + Δx, y5 + Δy), where (x5, y5) is the actual coordinate of the fifth frame. Connect the expected positions of all key points to form the expected limb contour. The expected limb contour is a closed polygon area that covers the predicted positions of the head, torso, and limbs.

[0084] When calculating the pixel-level displacement deviation between the actual segmentation edge of the current frame and the expected limb contour, compare the pixel coordinates of the expected limb contour and the actual segmentation edge one by one, and calculate the Euclidean distance between the expected position and the actual position of the same pixel point as the displacement deviation value.

[0085] If the displacement deviation value exceeds the dynamic displacement threshold, mark the area where the pixel point is located as an abnormal fragment. The dynamic displacement threshold is dynamically adjusted according to the average displacement deviation of the previous multiple frames. The specific adjustment method is as follows: calculate the average displacement deviation value of each key point in the previous multiple frames, and set the dynamic displacement threshold to a preset multiple range of the average displacement deviation value. This multiple range is set according to the actual scene requirements. For example, in a scene with large dynamic background interference, the multiple range is appropriately increased to tolerate higher deviations, while in a static background scene, the multiple range is reduced to improve the detection sensitivity.

[0086] In a scenario where the virtual background has dynamic special effects, such as a background containing rotating particles or gradient light effects, the displacement vectors of human body key points may exhibit abnormal fluctuations due to background interference. At this time, prediction based on the motion trajectories of the previous multiple frames can effectively distinguish real limb movements from background noise: If the displacement vector of a certain key point shows random fluctuations in the previous multiple frames (such as irregular directions and sudden speed changes), it is determined that the key point is affected by background interference, and its weight is reduced or it is directly excluded when predicting the motion trend of the current frame, thereby improving the accuracy of the expected limb contour.

[0087] In a scenario where the virtual background is a static solid color, such as a solid green screen, the displacement vectors of key points mainly reflect real limb movements. The deviation between the expected contour generated by linear extrapolation and the actual segmentation edge is small, and the dynamic displacement threshold adaptively decreases with the average deviation to avoid over-marking abnormal fragments.

[0088] After marking the abnormal fragments, morphological opening operation is performed on the area of the abnormal fragments. Morphological opening operation eliminates isolated noise points through the operations of erosion first and then dilation, while retaining the integrity of the continuous abnormal area. For example, isolated pixel blocks with an area smaller than the preset threshold are removed, and only the larger abnormal fragments adjacent to the human body area are retained. The marking results of the abnormal fragments are stored in the form of a binary mask image. In the binary mask image, the pixel values of the abnormal fragment area are 1, and the pixel values of the normal area are 0, which are used for targeted correction of the abnormal fragments in subsequent steps.

[0089] When generating multi-scale Gaussian heatmaps based on the coordinates of human body bone key points, small-scale high-weight Gaussian kernels are used to generate heatmaps for head and hand key points, and large-scale low-weight Gaussian kernels are used to generate heatmaps for torso key points. Specifically, for the coordinate positions of head key points (such as the top of the head and the chin) and hand key points (such as the fingertips and the wrists), heatmaps are generated using small-scale Gaussian kernels. The formula for the small-scale Gaussian kernel is: ; where represents the value of the small-scale Gaussian heatmap of the head and hand key points at the pixel coordinate ; represents the standard deviation of the small-scale Gaussian kernel, which controls the diffusion range of the heatmap. For example pixels; represents the two-dimensional coordinates of the head or hand key points, which are output by the human body bone key point detection algorithm; represents the coordinates of the current pixel, which is the pixel position to be calculated in the heatmap.

[0090] where is taken so that the coverage range of the heatmap does not exceed the preset radius (such as 5 pixels) around the key point.

[0091] The coordinate positions of the torso key points (such as shoulders and hips) generate heat maps through a large-scale Gaussian kernel, and the formula of the large-scale Gaussian kernel is expressed as: ; where represents the value of the large-scale Gaussian heat map of the torso key point at the pixel coordinate ; represents the standard deviation of the large-scale Gaussian kernel, satisfying , for example pixels, so that the coverage range of the heat map is expanded to a larger area around the key point (for example, 15 pixels); represents the two-dimensional coordinates of the torso key points (such as shoulders and hips), which are output by the human body bone key point detection algorithm.

[0092] When performing region adaptive fusion of the multi-scale Gaussian heat map and the initial human body segmentation result, in the head and hand regions, the heat map weight and the initial segmentation result are multiplied pixel by pixel to strengthen the edge gradient; specifically, the fusion formula for the head and hand regions is: ; where represents the confidence value of the fused segmentation map at the pixel coordinate ; represents the confidence value of the initial human body segmentation map at the pixel coordinate , with a range of [0,1]; represents the heat map enhancement coefficient, which controls the enhancement amplitude of the head and hand edges, for example .

[0093] In the torso region, directly retain the initial segmentation result, that is , to avoid over-modifying the torso contour.

[0094] After extracting the binary mask of the interference region according to the abnormal fragment marking result, a local region growth algorithm guided by the heat map weight is used to fill the holes in the interference region. The seed point selection rule of the region growth algorithm is: in the interference region, select the pixels with the heat map weight value greater than the preset threshold as the seed points, for example ; starting from the seed points, grow and expand to the surrounding neighborhoods (such as 8-neighborhood), and the growth condition is that the heat map weight value of the neighborhood pixels is not lower than γ times the weight value of the current pixel (for example, γ = 0.8), until the holes are filled.

[0095] Among them, represents the maximum weight value of all pixel values in the head and hand Gaussian heat maps, which is used to set the threshold benchmark for seed point selection in the region growth algorithm to ensure that only the region with the highest heat map weight is selected as the growth starting point.

[0096] When performing morphological closing operation with edge constraints on the non-interference region, a circular structuring element is used for dilation and erosion operations. The radius of the structuring element is rclose (for example, rclose = 3 pixels) to smooth the edges and maintain the contour integrity. Anisotropic diffusion filtering is performed on the interference region after region growing, and the diffusion coefficient is adjusted according to the heatmap weight value: ; where represents the anisotropic diffusion coefficient at the pixel coordinate , with a range of [0, 1]; represents the gradient magnitude of the fused segmentation map at the coordinate , indicating the degree of sharp change in pixel values; represents the gradient threshold parameter, which is used to adjust the sensitivity of the diffusion coefficient to gradient changes. For example .

[0097] The larger it is, the more details are retained in the regions with high heatmap weights, and the noise is smoothed in the regions with low weights.

[0098] In the scenario where the virtual background is a dynamic special effect, such as the background contains flashing light spots or floating particles, the holes in the interference region may be scattered at the edges of the head or hands. Region growing guided by the heatmap weight can preferentially fill from the high-weight regions (such as the fingertips) to ensure the continuity of the key part contours. In the static monochromatic background scenario, such as a solid green screen, the interference regions are less and concentrated. Anisotropic diffusion filtering can smooth the torso edges while retaining the details of the head.

[0099] When aligning the human skeleton trajectory of the corrected segmentation result with the synchronized English speech timestamps, time series sampling is performed on the human skeleton key point trajectory of the corrected segmentation result. The interval of time series sampling is consistent with the video frame rate. For example, when the video frame rate is 30 frames per second, the human skeleton key point coordinates are extracted every 33 milliseconds, generating the displacement time series data of the skeleton key points. The displacement time series data of the skeleton key points includes the two-dimensional coordinate sequences of the head, shoulders, elbows, and wrist key points.

[0100] Phoneme-level segmentation is performed on the synchronized English speech data. Phoneme-level segmentation locates the phoneme boundaries in the speech waveform through short-time energy detection and zero-crossing rate analysis, generating the peak timestamp sequence of the speech waveform. The peak timestamp sequence of the speech waveform marks the start and end time points of each phoneme. For example, when pronouncing the word "teamwork", the start timestamps of the phonemes / t / , / iː / , / m / , / w / , / ɜː / , / k / are 0ms, 120ms, 240ms, 360ms, 480ms, 600ms respectively.

[0101] When aligning the temporal data of skeletal key-point displacements with the peak timestamps of the speech waveform using the dynamic time warping (DTW) algorithm, the DTW algorithm constructs an accumulated cost matrix to find the optimal path between the skeletal displacement time series and the speech timestamps. The optimal path represents the non-linear mapping relationship between the two in the time dimension.

[0102] Specifically, the rate of change of the key-point velocity in the skeletal displacement time series (such as the hand movement speed) is used as the first-dimensional feature, and the peak energy of the speech waveform is used as the second-dimensional feature. The Euclidean distance between the two is calculated as the input to the cost matrix, and the minimum accumulated cost path is solved through dynamic programming to generate a skeletal-speech time mapping table. The skeletal-speech time mapping table records the time correspondence between the key frames of the skeletal actions and the speech phonemes. For example, the key frame of the hand waving upward (the 15th frame) is aligned with the starting phoneme / i / of the word "innovation" in the speech.

[0103] When extracting the semantic keywords that appear synchronously with the skeletal actions from the speech content based on the skeletal-speech time mapping table, according to the overlapping intervals between the key frames of the skeletal actions and the speech timestamps in the time mapping table, the speech text segments within the overlapping intervals are filtered. The speech text segments are converted into text content through a speech recognition engine.

[0104] For keyword extraction from the converted text content, the keyword extraction rules include: filtering the words associated with the corporate culture (such as "collaboration", "leadership"), and excluding the general meaningless words (such as "the", "and").

[0105] The extracted semantic keywords are bound and stored with the key frames of the skeletal actions. For example, the hand stretching action is bound with the speech segment "expand markets" to form a cross-modal associated data unit.

[0106] In a scenario where the virtual background has dynamic special effects, such as the background contains flashing light spots or floating particles, the temporal data of the skeletal key-point displacements may have noise fluctuations due to background interference. The DTW algorithm suppresses the misalignment caused by the noise fluctuations by adjusting the tolerance threshold of the path search.

[0107] For example, the tolerance threshold is set to 1.5 times the average speed of the skeletal displacement. If the speed fluctuation of a certain segment of the time series data exceeds this threshold, a penalty term is added to the cost matrix to reduce its impact on the optimal path. In a static monochromatic background scenario, such as a solid green screen, the noise of the skeletal displacement time series data is relatively low, and the DTW algorithm can directly adopt the standard path search strategy to improve the alignment efficiency.

[0108] The extraction results of voice and semantic keywords are stored in a structured data format, which includes three fields: timestamp, keyword text, and associated skeletal action type. For example, the timestamp "1200ms" corresponds to the keyword "strategic" and the associated action "cross hands in front of the chest".

[0109] The stored data is called in subsequent steps to screen for cross-cultural sensitive action segments, ensuring the spatio-temporal correlation between voice semantics and body movements.

[0110] When screening for culturally sensitive action segments based on the correlation between semantic keywords and the trajectory of human skeletal key points, calculate the time-space overlap degree between the semantic keywords and the trajectory of human skeletal key points. The definition of the time-space overlap degree is: within the voice time period when the semantic keyword appears, the product of the displacement amplitude of the corresponding human skeletal key points and the semantic intensity of the voice content. Among them, the displacement amplitude is obtained by calculating the change value of the Euclidean distance of the skeletal key points in three-dimensional space; the semantic intensity is obtained through a predefined corporate culture keyword weight table. For example, the keyword "teamwork" has a weight of 0.9, and "innovation" has a weight of 0.8. The product result is normalized to the range of 0 to 1 to generate a cross-modal correlation score.

[0111] When dynamically adjusting the cross-modal correlation score threshold according to the cultural type of the target enterprise, if the target enterprise is an American enterprise, the score threshold is set to 0.7; if it is a Japanese enterprise, the score threshold is reduced to 0.6 to adapt to the differences in the sensitivity of body movements in different cultures.

[0112] Based on the analysis of historical enterprise interview data, the American enterprise culture has a lower sensitivity to body movements (such as allowing large gestures), and the cross-modal correlation score threshold is set to 0.7; due to the introverted cultural characteristics of Japanese enterprises, the sensitivity needs to be reduced, and the threshold is adjusted to 0.6. The threshold is optimized through testing 1000 groups of sample data with virtual backgrounds.

[0113] Screen for candidate action segments with a cross-modal correlation score exceeding the threshold. For example, a hand stretching action segment with a cross-modal correlation score of 0.75 (in the American enterprise scenario) is retained as a candidate segment, and a nodding action segment with a score of 0.65 (in the Japanese enterprise scenario) is also retained due to the reduced threshold.

[0114] When verifying the cultural sensitivity of candidate action segments, conflict detection is performed based on the predefined gesture amplitude threshold range and sitting posture uprightness score gradient in the cross-cultural body language rule library.

[0115] The gesture amplitude threshold range is obtained through data analysis and statistics of historical corporate interview videos. For example, American companies allow a gesture amplitude range of 50 - 150 pixels, while Japanese companies limit it to 30 - 100 pixels; the sitting posture uprightness scoring gradient is calculated based on the angle between the torso key points and the vertical axis. For example, a score of excellent is given when the angle is less than 10 degrees, good when it is 10 - 20 degrees, and unqualified when it exceeds 20 degrees.

[0116] If the gesture amplitude or sitting posture score of a candidate action segment exceeds the range allowed by the rule library, it is determined as a conflicting action and filtered. For example, a hand waving action with an amplitude of 160 pixels is filtered in a Japanese corporate scenario.

[0117] When integrating the verified action segments, the culturally sensitive action segments that meet the gesture amplitude threshold and have a qualified sitting posture score are merged in chronological order to generate a set of culturally sensitive action segments. Each segment in the set of culturally sensitive action segments contains a start timestamp, an end timestamp, associated semantic keywords, and an action type label. For example, the timestamp "1200ms - 1500ms" is associated with the keyword "strategic deployment" and the action type "crossed hands in front of the chest", which is used for subsequent steps to generate suggestions for optimizing body language.

[0118] In a dynamic virtual background scenario, such as a background containing rotating particle effects, the candidate action segments may have calculation errors in displacement amplitude due to background interference. At this time, noise is suppressed by adding smoothing filtering for displacement amplitude (such as the moving average method) to ensure the accuracy of the cross-modal correlation score. In a static single-color background scenario, such as a solid green screen, the calculation error of displacement amplitude is relatively low, and the original data is directly used for scoring.

[0119] Dynamic virtual backgrounds (such as particle effects) are prone to causing bone trajectory noise. The American company threshold of 0.7 is set to be more sensitive to accommodate large-scale actions allowed by the culture, while the Japanese company threshold of 0.6 is set to strictly filter to adapt to the conservative cultural characteristics.

[0120] The gesture amplitude threshold and sitting posture scoring gradient in the cross-cultural body language rule library are obtained through data analysis and statistics of historical corporate interview videos. For example, by analyzing the gesture amplitude distribution of successful candidates in 1000 American corporate interview videos, the 5% to 95% percentile values are taken as the threshold range. The rule library supports custom updates according to corporate needs. For example, when adding German corporate culture rules, it is necessary to import the gesture amplitude sample data of the corresponding culture and recalculate the threshold.

[0121] Obtain the action type tags and corresponding timestamps in the culturally sensitive action segment set. The action type tags include gesture action categories (such as "crossing hands in front of the chest", "waving hands") and sitting posture categories (such as "leaning forward", "leaning backward"). The timestamps record the start and end time points of the action segment. Query the gesture amplitude threshold range and sitting posture uprightness scoring gradient in the cross-cultural body language rule library according to the action type tags. The cross-cultural body language rule library stores the amplitude limits and sitting posture standards of various actions under different cultural types. For example, the amplitude threshold of "waving hands" in the culture of American enterprises is 50-150 pixels, and in the culture of Japanese enterprises, it is 30-100 pixels; the sitting posture uprightness scoring gradient is divided based on the angle between the torso key points and the vertical axis. For example, when the angle is less than 10 degrees, the score is grade A (excellent), 10-20 degrees is grade B (good), and more than 20 degrees is grade C (unqualified).

[0122] When comparing and analyzing the actual gesture amplitude in the culturally sensitive action segment with the gesture amplitude threshold range, extract the maximum displacement distance of the hand key points in the action segment as the actual gesture amplitude. If the actual gesture amplitude exceeds the threshold range of the corresponding action type in the rule library, it is marked as an action to be corrected. For example, in the scenario of American enterprises, the amplitude of "waving hands" in a certain segment is 160 pixels (threshold 50-150 pixels), which is marked as an action to be corrected; in the scenario of Japanese enterprises, if the amplitude of the same action is 110 pixels (threshold 30-100 pixels), it is also marked as an action to be corrected.

[0123] When generating body language suggestions adapted to the target enterprise scenario based on the marking results of the actions to be corrected and the sitting posture uprightness scoring gradient, calculate the gesture amplitude adjustment range for the actions to be corrected. The adjustment range is set as the ±10% tolerance interval of the rule library threshold. For example, the suggestion for "waving hands" in American enterprises is adjusted to 50-150 pixels, and after tolerance, it is 45-165 pixels; the sitting posture optimization angle is deduced inversely according to the scoring gradient. For example, when the sitting posture score is grade C (angle 20 degrees), it is recommended to adjust to less than 10 degrees. The generated body language suggestions include text descriptions and visual examples. For example, "The amplitude of waving hands needs to be controlled within the range of 45-165 pixels, and the recommended forward leaning angle of the sitting posture is less than 10 degrees", and a standard action schematic diagram is attached.

[0124] In a dynamic virtual background scenario, such as a background containing flashing light spots or floating particles, the actual gesture amplitude may be overestimated due to background interference. At this time, reduce the noise impact by adding smoothing filtering processing for amplitude calculation (such as the moving average method) to ensure the accuracy of the comparison with the threshold. In a static single-color background scenario, such as a pure green screen, directly use the original amplitude data for comparative analysis. The data in the cross-cultural body language rule library is obtained through the analysis and statistics of historical enterprise interview videos. For example, analyze the gesture amplitude distribution of successful candidates in 500 American enterprise interviews, and take the 5% to 95% percentile values as the threshold range, and dynamically update the threshold regularly according to the new data.

[0125] The output format of the body language suggestions includes a list of correction instructions associated with timestamps and a visual timeline. For example, during video playback, the timeline marks the segments of actions to be corrected and overlays text prompts, and users can click to view detailed suggestions. The list of correction instructions is classified and summarized by action type. For example, it summarizes all the "hand waving" over-limit segments and their recommended adjustment ranges, facilitating targeted improvement by users.

[0126] Embodiment 2: Figure 2 A structural schematic diagram of the AI interview English body language assisted teaching system integrating computer vision according to the present invention is given. The AI interview English body language assisted teaching system integrating computer vision includes a virtual video acquisition module, an image segmentation generation module, a motion trend prediction module, a contour correction processing module, a voice-skeleton alignment module, a cultural action screening module, and a body language suggestion generation module.

[0127] Virtual video acquisition module: Obtain the interview video stream under the virtual background through a camera, and extract the RGB image and background mask image of the video frame.

[0128] Image segmentation generation module: Input the RGB image and background mask image into the segmentation network to generate the initial human body segmentation result of the current frame.

[0129] Motion trend prediction module: Predict the limb motion trend of the current frame based on the initial human body segmentation results of the previous multiple frames, compare the displacement deviation between the predicted contour and the actual segmentation edge, and mark those exceeding the dynamic threshold as abnormal fragments.

[0130] Contour correction processing module: Strengthen the head and hand edge contours by combining the weights of the human key point heat maps, and remove abnormal fragments to obtain the corrected segmentation result.

[0131] Voice-skeleton alignment module: Align the human skeleton trajectory of the corrected segmentation result with the synchronous English speech timestamps, and extract the semantic keyword of the speech.

[0132] Cultural action screening module: Screen the culturally sensitive action segments according to the relevance between the semantic keywords and the skeleton trajectory.

[0133] Body language suggestion generation module: Match the culturally sensitive action segments with the gesture amplitude and sitting posture standards in the cross-cultural rule library to generate body language suggestions adapted to the enterprise scenario.

[0134] The above formulas are all dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and threshold selection in the formula are set by those skilled in the art according to the actual situation.

[0135] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0136] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0137] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0138] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0139] In addition, in each embodiment of the present application, each functional module can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0140] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0141] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0142] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. AI interview English body language assisted teaching method integrating computer vision, characterized by: The steps include: S1. Obtain the interview video stream under the virtual background through the camera, and extract the RGB image and background mask image of the video frame; S2, input the RGB image and the background mask image into the segmentation network to generate the initial human segmentation result of the current frame; S3, predicting the limb movement trend of the current frame based on the initial human body segmentation results of the previous multiple frames, comparing the displacement deviation between the predicted contour and the actual segmentation edge, and marking the abnormal fragments with excessive dynamic thresholds; S4, combining the weights of the key points of the human body heat map to strengthen the edge contours of the head and hands, and remove abnormal fragments to obtain the corrected segmentation results; S5, aligning the human skeleton trajectory of the corrected segmentation result with the synchronized English speech timestamp, and extracting speech semantic keywords; S6, screening culturally sensitive action clips based on the correlation between semantic keywords and skeletal trajectories; S7. Match culturally sensitive action clips with the gesture amplitude and sitting posture standards of the cross-cultural rule library to generate body language suggestions suitable for corporate scenarios.

2. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 1 is characterized in that: The interview video stream under the virtual background is obtained through the camera, and the RGB image and background mask of the video frame are extracted, including: S1a, collect the interview video stream containing the virtual background in real time through the camera; S1b, decode the interview video stream frame by frame and extract the RGB image of each video frame; S1c, performing color key matting processing on the RGB image to generate a binary background mask image corresponding to the virtual background area; S1d, store the RGB image and the corresponding background mask image into the video frame buffer queue.

3. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 1 is characterized in that: Input the RGB image and background mask image into the segmentation network to generate the initial human segmentation result of the current frame, including: S2a, inputting the RGB image and the background mask image into the segmentation network, extracting the color distribution characteristics of the background mask image, the color distribution characteristics including the color level mean and dispersion of the main color channel; S2b, classifying the virtual background type into a static single-color background or a dynamic multi-color background based on the color level mean and dispersion of the main color channel; S2c, dynamically adjust the color gamut difference threshold according to the virtual background type: if it is a static single-color background, the color gamut difference threshold is set to a fixed range; if it is a dynamic multi-color background, the color gamut difference threshold is linearly expanded as the discreteness of the main color channel increases; S2d, performing pixel-level filtering on the RGB image based on the adjusted color gamut difference threshold to generate an initial human segmentation result of the current frame.

4. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 1 is characterized in that: The limb movement trend of the current frame is predicted based on the initial human body segmentation results of the previous multiple frames. The displacement deviation between the predicted contour and the actual segmentation edge is compared, and the over-dynamic threshold is marked as abnormal fragments, including: S3a, obtaining the initial human body segmentation results of the previous multiple frames, extracting the displacement vectors of the key points of the human body in each frame, and generating the key point motion trajectory; S3b, predicting the limb movement trend of the current frame based on the key point motion trajectory, and generating the expected limb contour; S3c, calculating the pixel-level displacement deviation between the actual segmentation edge of the current frame and the expected limb contour, and if the pixel-level displacement deviation exceeds the dynamic displacement threshold, marking the corresponding area as an abnormal fragment; S3d, the dynamic displacement threshold is dynamically adjusted according to the average displacement deviation of the previous multiple frames.

5. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 1 is characterized in that: Combine the weights of the key points of the human body heat map to strengthen the edge contours of the head and hands, and remove abnormal fragments to obtain the corrected segmentation results, including: S4a, generating a multi-scale Gaussian heat map based on the coordinates of the key points of the human skeleton, using a small-scale high-weight Gaussian kernel for the key points of the head and hands, and a large-scale low-weight Gaussian kernel for the key points of the torso; S4b, regional adaptive fusion of the multi-scale Gaussian heat map and the initial human segmentation result: the edge gradient of the head and hand regions is enhanced according to the heat map weight, and the original contour of the torso region is retained according to the initial segmentation result; S4c, extracting the binary mask of the interference area according to the abnormal debris marking result, and filling the holes in the interference area using the local region growing algorithm guided by the heat map weight; S4d, performing edge-constrained morphological closing operations on the non-interference regions, performing anisotropic diffusion filtering on the interference regions after region growth, and generating a modified segmentation result.

6. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 5 is characterized in that: Anisotropic diffusion filtering is performed on the interference area after region growing, and the diffusion coefficient is adjusted according to the weight value of the heat map: ;in, Represented in pixel coordinates The anisotropic diffusion coefficient at , Indicates that the fused segmentation map is in coordinate The gradient amplitude at Represents the gradient threshold parameter.

7. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 1 is characterized in that: Align the human skeleton trajectory of the corrected segmentation result with the synchronized English speech timestamp and extract speech semantic keywords, including: S5a, performing time series sampling on the trajectory of the human skeleton key points of the corrected segmentation result to generate time series data of the displacement of the skeleton key points; S5b, performing phoneme-level segmentation on the synchronized English speech data to generate a speech waveform peak time stamp sequence; S5c, aligning the skeleton key point displacement time series data with the speech waveform peak timestamp sequence through a dynamic time warping algorithm, and establishing a skeleton-speech time mapping table; S5d. Extracting semantic keywords that appear synchronously with the skeleton movements in the speech content based on the skeleton-speech time mapping table, wherein the semantic keywords include target corporate culture-related vocabulary.

8. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 1 is characterized in that: Filter culturally sensitive action clips based on the correlation between semantic keywords and skeletal trajectories, including: S6a, calculating the time-space overlap between the semantic keywords and the trajectories of the key points of the human skeleton, and generating a cross-modal relevance score, where the time-space overlap is the product of the displacement amplitude of the corresponding key points of the human skeleton during the time period when the semantic keywords appear and the semantic intensity of the speech content; S6b, dynamically adjusting the cross-modal relevance score threshold according to the culture type of the target enterprise, and screening candidate action clips whose cross-modal relevance scores exceed the threshold; S6c, verify the cultural sensitivity of candidate action clips, and filter out action clips that conflict with the target corporate culture based on the gesture amplitude threshold range and sitting posture correctness score gradient predefined in the cross-cultural body language rule library; S6d. Integrate the verified action clips to generate a set of culturally sensitive action clips.

9. The AI ​​interview English body language assisted teaching method integrating computer vision according to claim 1 is characterized in that: Match culturally sensitive action clips with the gesture amplitude and sitting posture standards of the cross-cultural rule library to generate body language suggestions that are suitable for corporate scenarios, including: S7a, obtaining action type labels and corresponding timestamps in a set of culturally sensitive action clips; S7b, querying the gesture amplitude threshold range and sitting posture correctness score gradient in the cross-cultural body language rule library according to the action type label; S7c, comparing and analyzing the actual gesture amplitude in the culturally sensitive action clip with the gesture amplitude threshold range, and marking it as an action to be corrected if the actual gesture amplitude exceeds the threshold range; S7d. Generate body language suggestions adapted to the target enterprise scenario based on the marking results of the action to be corrected and the sitting posture correctness score gradient. The suggestions include the gesture amplitude adjustment range and the sitting posture optimization angle.

10. An AI interview English body language assisted teaching system integrating computer vision, used to implement the AI ​​interview English body language assisted teaching method integrating computer vision as described in any one of claims 1 to 9, characterized in that: It includes virtual video acquisition module, image segmentation generation module, motion trend prediction module, contour correction processing module, voice skeleton alignment module, cultural action screening module and limb suggestion generation module; Virtual video acquisition module: obtains the interview video stream under the virtual background through the camera, and extracts the RGB image and background mask image of the video frame; Image segmentation generation module: input the RGB image and background mask image into the segmentation network to generate the initial human segmentation result of the current frame; Motion trend prediction module: predicts the limb motion trend of the current frame based on the initial human body segmentation results of the previous multiple frames, compares the displacement deviation between the predicted contour and the actual segmentation edge, and marks the over-dynamic threshold as abnormal fragments; Contour correction processing module: Combine the weights of the key points of the human body heat map to strengthen the edge contours of the head and hands, and remove abnormal fragments to obtain the corrected segmentation results; Speech skeleton alignment module: aligns the human skeleton trajectory of the corrected segmentation result with the synchronized English speech timestamp to extract speech semantic keywords; Cultural action screening module: Screen culturally sensitive action clips based on the correlation between semantic keywords and skeleton trajectories; Body language suggestion generation module: matches culturally sensitive action clips with the gesture amplitude and sitting posture standards of the cross-cultural rule library to generate body language suggestions suitable for corporate scenarios.

Citation Information

Cited By

  • Enterprise operation data real-time analysis system based on mobile BI

    CN120448842A

  • Limb movement learning clue labeling method and system based on cross-modal mapping

    CN120977153A