Machine tool processing vibration state recognition method and system based on video analysis

CN122618534BActive Publication Date: 2026-09-25CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611047341.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-25
Estimated Expiration
2046-07-15

AI Technical Summary

Technical Problem

[0003]现有机床加工振动状态识别方案利用加工过程图像、加工后表面纹理图像或者视频数据识别加工振动状态,例如通过相机采集加工区域图像,提取刀具、工件或加工表面纹理的变化特征,再根据图像中的边缘偏移、纹理周期、画面抖动或局部运动变化推断加工振动状态,以上方案能够以非接触方式获取加工区域信息,减少在机床结构上额外布置振动传感器的需求,并且能够将加工过程中的可见运动变化转化为图像数据进行分析;但在立式加工中心铣削薄壁框类零件内侧壁的过程中,以上方案的振动识别方法通常会将画面中目标轮廓消失现象直接作为遮挡事件处理,并据此继续进行轨迹补偿或异常片段提取;但在刀具周期性经过薄壁内侧壁前方的具体加工状态下,画面中的轮廓消失并不必然由刀具遮挡造成,还可能由切屑飞溅、冷却液反光、刀具阴影、局部模糊、非目标轮廓交叠或背景边界干扰引起,例如,在刀具尚未真正经过待分析内侧壁边缘邻近位置时,视频画面中的某段边缘可能因切屑遮挡而短时消失,若仅依据边缘消失后又出现便将该时间段作为遮挡区间并进入轨迹补全,则会把非刀具遮挡造成的图像缺失误作为薄壁边缘振动轨迹的一部分,使得后续补全得到的轨迹会落在错误对象或错误区间上,最终使异常振动时间窗的起止位置和对象归属发生偏差

Benefits of technology

1、本方案在加工视频数据中逐帧追踪薄壁零件候选边缘对象,可形成同一边缘对象的帧间连续位置关系,从而避免将不同边缘或背景边缘误作为同一分析对象;识别其依次呈现可见过程、缺失过程和再显过程,可确定缺失并非目标消失而是具有前后承接的中断区间,从而为遮挡判定提供可恢复的对象依据;将缺失过程的时间区间与数控加工执行数据对应的时序范围比对,并仅在完全落入该时序范围内时确立目标遮挡区间,可排除非加工时段、图像噪声或非刀具因素造成的缺失,使目标遮挡区间同时具备视觉连续性依据和加工执行时序依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618534B_ABST
    Figure CN122618534B_ABST
Patent Text Reader

Abstract

The application discloses a machine tool machining vibration state recognition method and system based on video analysis, relates to the technical field of image data processing and video target recognition, and comprises the following steps: acquiring machining video data and numerical control machining execution data of a target object; performing space-time alignment on the machining video data and the numerical control machining execution data, determining a corresponding dynamic adjacent processing area of a tool space movement process in a video picture, and extracting a thin-walled part candidate edge object with direction continuity in the dynamic adjacent processing area. According to the scheme, the thin-walled part candidate edge object is tracked frame by frame in the machining video data, the interframe continuous position relationship of the same edge object can be formed, different edges or background edges are avoided from being mistaken as the same analysis object, the visible process, the missing process and the reappearing process are recognized in sequence, and it can be determined that the missing is not the disappearance of the target but an interruption interval with a front and back connection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image data processing and video target recognition technology, specifically to a method and system for identifying machine tool processing vibration state based on video analysis. Background Technology

[0002] Vibration identification during machine tool processing is an important technical aspect of machining process monitoring. Especially in material removal processes such as milling, boring, and turning, the interaction between the tool and the workpiece can cause minute vibrations in the workpiece, tool, or spindle system. When the vibration is within the normal machining fluctuation range, its impact on the machining process is relatively limited. However, when the vibration amplitude, duration, or state changes abnormally with the machining process, it can easily lead to problems such as surface ripples, dimensional errors, accelerated tool wear, or even machining instability.

[0003] Existing machine tool vibration state recognition schemes utilize images of the machining process, post-machining surface texture images, or video data to identify the vibration state. For example, images of the machining area are captured by a camera, and the variation features of the tool, workpiece, or machined surface texture are extracted. Then, the machining vibration state is inferred based on edge offset, texture periodicity, image jitter, or local motion changes in the image. These schemes can acquire machining area information in a non-contact manner, reducing the need for additional vibration sensors on the machine tool structure, and can convert visible motion changes during machining into image data for analysis. However, during the milling of the inner wall of thin-walled frame-like parts in a vertical machining center, the vibration recognition methods of these schemes typically treat the disappearance of the target outline in the image as an occlusion event and continue the trajectory analysis accordingly. Trajectory compensation or abnormal segment extraction; however, in the specific machining state where the tool periodically passes in front of the thin-walled inner sidewall, the disappearance of the outline in the picture is not necessarily caused by tool occlusion. It may also be caused by chip splash, coolant reflection, tool shadow, local blur, non-target outline overlap or background boundary interference. For example, before the tool has actually passed the position near the edge of the inner sidewall to be analyzed, a certain edge in the video picture may disappear briefly due to chip occlusion. If the time period after the edge disappears and reappears is taken as the occlusion interval and entered into trajectory completion, the image loss caused by non-tool occlusion will be mistakenly taken as part of the thin-walled edge vibration trajectory, so that the trajectory obtained by subsequent completion will fall on the wrong object or wrong interval, and ultimately the start and end positions of the abnormal vibration time window and the object attribution will be deviated. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for identifying machine tool processing vibration status based on video analysis, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the technical solution of the present invention is as follows: In a first aspect, the present invention discloses a machine tool machining vibration state recognition method based on video analysis, applied to the identification of abnormal time windows during the milling vibration of the inner wall of thin-walled frame parts in a vertical machining center, comprising the following steps: Acquire the processing video data and CNC machining execution data of the target object; The machining video data and the CNC machining execution data are spatiotemporally aligned to determine the dynamic proximity processing area corresponding to the tool spatial movement process in the video frame, and candidate edge objects of thin-walled parts with directional continuity are extracted within the dynamic proximity processing area. In the processing video data, the candidate edge objects of the thin-walled part are tracked frame by frame. When the candidate edge objects of the thin-walled part are identified to show a temporal state change of a visible process, a missing process and a re-enactment process in sequence, and when the time interval of the missing process falls completely within the temporal range corresponding to the CNC machining execution data, the time interval corresponding to the missing process is established as the target occlusion interval. Extract the candidate edge objects of the thin-walled parts at the front and rear ends of the target occlusion area and construct the candidate completion trajectory connecting the two ends accordingly; The candidate completion trajectory is back-end verified based on the candidate edge object of the thin-walled part that is re-displayed, and the back-end verification result, the target occlusion region and the candidate edge object of the thin-walled part are aggregated to generate an anomaly recognition result.

[0006] Secondly, this invention discloses a machine tool processing vibration state recognition system based on video analysis, comprising: The data acquisition module is used to acquire the processing video data and CNC machining execution data of the target object; The tool proximity edge filtering module is used to perform spatiotemporal alignment between the machining video data and the CNC machining execution data, determine the dynamic proximity processing area corresponding to the tool spatial movement process in the video frame, and extract thin-walled part candidate edge objects with directional continuity within the dynamic proximity processing area. The occlusion interval determination module is used to track the candidate edge objects of the thin-walled part frame by frame in the processing video data, and when the candidate edge objects of the thin-walled part are identified to show the temporal state changes of the visible process, the missing process and the re-display process in sequence, and when the time interval of the missing process falls completely within the temporal range corresponding to the CNC machining execution data, the time interval corresponding to the missing process is established as the target occlusion interval. The occlusion trajectory completion module is used to extract the candidate edge objects of the thin-walled parts at the front and rear ends of the target occlusion area and construct the candidate completion trajectory connecting the two ends accordingly. The anomaly detection result output module is used to perform backend verification on the candidate completion trajectory based on the candidate edge object of the thin-walled part that is re-displayed, and to aggregate the backend verification result, the target occlusion region and the candidate edge object of the thin-walled part to generate anomaly detection result.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This solution tracks candidate edge objects of thin-walled parts frame by frame in the processing video data, which can form a continuous positional relationship between the same edge object in frames, thereby avoiding mistaking different edges or background edges as the same analysis object; it identifies the visible process, missing process and re-appearance process in sequence, which can determine that the missing is not the disappearance of the target but an interrupted interval with preceding and following connections, thus providing a recoverable object basis for occlusion judgment; it compares the time interval of the missing process with the time sequence range corresponding to the CNC machining execution data, and establishes the target occlusion interval only when it completely falls within the time sequence range, which can exclude the missing caused by non-processing periods, image noise or non-tool factors, so that the target occlusion interval has both visual continuity basis and machining execution time sequence basis.

[0008] 2. This scheme generates a tolerance frame length by dividing the parsing frame rate by the CNC timing step size, ensuring that the video judgment window is consistent with the CNC execution rhythm and reducing state misjudgments caused by asynchronous sampling; it generates a dynamic matching baseline by extracting and normalizing the initial pixel structural features, providing a common reference for subsequent similarity comparisons and suppressing the influence of illumination and scale fluctuations; it constructs a temporal sliding window based on the tolerance frame length and calculates structural similarity frame by frame, ensuring that visibility, absence, and re-display are all confirmed by statistical analysis of consecutive frames, avoiding misidentification caused by single-frame noise and improving the stability and accuracy of occlusion interval determination. Attached Figure Description

[0009] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein: Figure 1 A flowchart illustrating the steps of the machine tool processing vibration state identification method based on video analysis provided by the present invention; Figure 2 A schematic diagram of the process for generating a dynamic proximity processing region provided by the present invention; Figure 3 This is a schematic diagram of the process for generating candidate edge objects for thin-walled parts provided by the present invention; Figure 4 This is a schematic diagram of the process for generating backend verification results provided by the present invention; Figure 5 A schematic diagram of the module functions of the machine tool processing vibration state recognition system based on video analysis provided by the present invention. Detailed Implementation

[0010] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.

[0011] Application Overview: Existing solutions for identifying machine tool vibration states using machining process images or video data typically infer vibration states based on edge offsets, texture periods, image jitter, or local motion changes in the tool, workpiece, or machined surface texture. When milling the inner wall of a thin-walled frame-like part in a vertical machining center, the tool periodically passes in front of the inner wall, and the target outline in the video may briefly disappear and reappear. However, this disappearance does not necessarily correspond to tool occlusion; it could also be caused by chip splashes, coolant reflections, tool shadows, local blurring, overlapping of non-target outlines, or background boundary interference. If existing processing chains simply use the time interval between edge disappearance and reappearance as the occlusion interval and continue with trajectory compensation or abnormal segment extraction, there is a lack of corresponding constraints between the formation of the occlusion interval and the actual position of the tool passing near the edge of the inner wall being analyzed. This leads to a risk of deviation between the object and the interval basis upon which the abnormal time window identification relies.

[0012] When milling the inner wall of a thin-walled frame-like part in a vertical machining center, a camera captures video of the machining area. The object to be analyzed is the edge of the thin-walled inner wall in the video frame. During machining, before the tool has actually passed the vicinity of this inner wall edge, a section of the edge may briefly disappear from the image due to chip splashes or coolant reflections, and then reappear. If existing vibration recognition solutions directly identify the corresponding time period as an occlusion interval and proceed to subsequent trajectory compensation based solely on the change in the image from visible to disappearing and then reappearing of this edge, then this time period actually exposes image loss issues not caused by tool occlusion, rather than target contour occlusion caused by the tool's spatial movement. The resulting processed object will deviate from the actual edge of the thin-walled part being analyzed, and the front and rear contours on which the trajectory compensation is based may correspond to incorrect objects or incorrect intervals.

[0013] If the aforementioned issues are not addressed, the temporary disappearance of contours in the video data will be propagated along the existing processing chain as a deviation in occlusion region confirmation, thus causing subsequent trajectory compensation to be based on image gaps not caused by tool occlusion. This deviation will cause the compensation trajectory to fall into the wrong object or the wrong time interval, and will continue to affect the abnormal segment extraction process, causing the start and end positions of the abnormal vibration time window to shift, or causing a mismatch in the attribution relationship between the abnormal vibration time window and the candidate edge objects of the thin-walled part. Ultimately, the vibration state recognition result based on this abnormal time window will have problems such as unclear object attribution, inaccurate time boundaries, and abnormal processing chain, affecting the correspondence between the machining process monitoring results and the actual milling process.

[0014] After introducing the basic concept of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0015] Example 1: This embodiment provides a machine tool machining vibration state recognition method based on video analysis. The method is applied to the identification of abnormal time windows during the milling vibration of the inner wall of thin-walled frame parts in a vertical machining center. The method can be executed by a machining process video analysis device, which includes at least an industrial camera mounted on the protective door or one side of the spindle box of the vertical machining center, covering the inner wall of the thin-walled frame part to be machined, a data acquisition interface communicating with the CNC system, and an edge computing unit carrying video parsing and timing determination logic. For ease of understanding, the following description uses the case of climb milling the inner wall of an aluminum alloy thin-walled frame part on a vertical machining center as the operating scenario, but this does not constitute a limitation on the application scenario. The method is also applicable to titanium alloy thin-walled cavities, magnesium alloy frames, thin-walled cylindrical parts, and other milling conditions where the image is easily lost temporarily due to the tool periodically passing in front of the edge to be analyzed during material removal.

[0016] Thin-walled frame parts in vertical machining centers refer to frame-shaped structural components whose wall thickness is significantly smaller than their outline dimensions, and whose inner walls are prone to out-of-plane micro-vibrations under milling forces. The inner wall to be machined typically appears as a slender, continuous edge in the video feed. Because thin-walled frame parts are relatively weak, the micro-vibrations caused by the interaction between the tool and the workpiece during milling of the inner wall are directly reflected as positional jitter of this slender edge in the frame. Therefore, the frame-by-frame positional evolution of this edge is an observable object carrying vibration information. This is the practical basis for choosing the edge object in the video feed, rather than the contact sensor signal on the machine tool structure, as the analysis carrier in this embodiment.

[0017] The initial data processed in this embodiment includes two types.

[0018] Acquire the processing video data and CNC machining execution data of the target object; The machining video data consists of a sequence of images continuously captured by an industrial camera during the milling of the inner wall, each frame bearing a timestamp. The CNC machining execution data is a real-time output record from the CNC system during the same machining process, also bearing an execution timestamp. This record includes at least the tool's three-dimensional coordinates over time in the machine tool coordinate system, the tool feed rate, and the spindle speed. The machining video data is acquired from the industrial camera, allowing vibration information to be collected non-contactly. This eliminates the need for additional contact vibration sensors on thin-walled frame parts or the spindle system, preventing the sensor's added mass from altering the vibration characteristics of the thin-walled structure. The CNC machining execution data is read directly from the CNC system, providing a definite temporal basis for the tool's actual spatial movement. This allows for temporal interpretation of image phenomena in the machining video data, rather than relying solely on brightness or edge changes in the image to infer whether the tool actually passed the adjacent position of the inner wall to be analyzed. Both types of initial data carry independent timestamps, providing a time reference for subsequent spatiotemporal alignment. This is a prerequisite for all subsequent decision-making logic in this embodiment to unfold on a unified timeline.

[0019] After acquiring the processing video data and CNC machining execution data, such as Figure 1 As shown, this embodiment unfolds step by step in the order of tool-adjacent edge screening, occlusion interval determination, occlusion trajectory completion, and anomaly identification result output. Each step is interconnected through clear data flow, forming a complete judgment chain around the candidate edge objects of thin-walled parts.

[0020] Spatiotemporal alignment of machining video data and CNC machining execution data is performed to determine the dynamic proximity processing area corresponding to the tool spatial movement process in the video frame, and candidate edge objects of thin-walled parts with directional continuity are extracted within the dynamic proximity processing area. It should be noted that the purpose of this step is to first align the machining video data and CNC machining execution data on the same timeline and in the same frame coordinate system. Based on this, the actual adjacent area occupied by the tool in the frame is delineated, and edges that truly belong to the inner wall to be machined and are directionally consistent between adjacent frames are selected as the subsequent analysis objects. Spatiotemporal alignment allows the 3D coordinates of the tool recorded by the CNC system to be projected onto the frame position of the machining video data. The dynamic proximity processing area constrains the spatial range of subsequent judgments to the local area where the tool actually operates. Candidate edge objects for thin-walled parts with directional continuity narrow the analysis objects from any visible edge in the frame to a single, continuous, frame-by-frame traceable edge of the thin-walled inner wall. The specific methods for determining the dynamic proximity processing area and extracting candidate edge objects for thin-walled parts will be elaborated in subsequent sections; here, only its object selection responsibility in the overall process is explained.

[0021] This feature addresses the problem in the prior art where the disappearance of any outline in the image is treated as an occlusion event without distinction. This is because existing solutions do not differentiate the spatial location and object affiliation of missing edges in the image, easily including edge missingness caused by chip splashes, coolant reflections, tool shadows, local blurring, overlapping of non-target outlines, or background boundary interference in subsequent processing. In this embodiment, before entering the occlusion determination, the spatial range of the analysis is constrained by the dynamic proximity processing area corresponding to the tool's spatial movement process, and the homology of the analyzed objects is constrained by the directional continuity. This ensures that only edges that truly belong to the inner wall to be processed and are within the actual operating area of ​​the tool become the subsequent determination objects. Therefore, this is not something that can be naturally obtained by adjusting the edge detection sensitivity in existing solutions that use the changes of the entire image edge as input.

[0022] In the video processing data, candidate edge objects of thin-walled parts are tracked frame by frame. When the candidate edge objects of thin-walled parts are identified to show the temporal state changes of visible process, missing process and re-showing process in sequence, and when the time interval of the missing process falls completely within the temporal range corresponding to the CNC machining execution data, the time interval corresponding to the missing process is established as the target occlusion interval. It should be noted that in this step, the presence of thin-walled part candidate edge objects in the image is tracked frame by frame along the timeline. When the thin-walled part candidate edge object shows three temporal state changes in sequence—first visible, then missing, and then reappearing—the time interval corresponding to the missing part is not immediately considered as tool occlusion. Only when the time interval of the missing process falls completely within the time range corresponding to the CNC machining execution data, and the tool has indeed passed through the vicinity of the inner wall to be machined, is the time interval corresponding to the missing process established as the target occlusion interval. The target occlusion interval refers to the time interval that, after being endorsed by the CNC machining execution data time sequence, is identified as the time interval corresponding to the image missing caused by the tool actually passing through the vicinity of the inner wall to be analyzed. It is the internal anchor point for subsequent trajectory completion and anomaly identification. The specific identification methods for the three temporal states of the visible process, the missing process, and the reappearing process will be elaborated in the corresponding sections later.

[0023] This feature addresses the problem in the prior art where the time period after an edge disappears and reappears is used as an occlusion interval for trajectory completion, thus mistaking image loss caused by non-tool occlusion as part of the thin-walled edge vibration trajectory. This is because before the tool actually passes the vicinity of the inner wall edge to be analyzed, a certain edge in the image may disappear briefly due to chip occlusion, reflection, or local blurring. If this is used as an occlusion interval without constraint, both the object and the interval to be completed will be offset. This embodiment uses whether the time interval of the missing process completely falls within the corresponding time sequence range of the CNC machining execution data as a necessary condition for occlusion to be established. The actual movement time sequence of the tool is used to endorse the image missing event, so that the timing of occlusion determination is no longer determined by the image itself, but by whether the tool has actually passed through. Therefore, it cannot be equivalently replaced by existing solutions that only rely on changes in the presence of image edges by increasing the frame threshold for edge disappearance determination.

[0024] Extract candidate edge objects of thin-walled parts at the front and rear ends of the target occlusion area and construct candidate completion trajectories connecting the two ends accordingly; It should be noted that, since the candidate edge objects of the thin-walled parts are missing in the image within the target occlusion area, the edge positions within this area are difficult to observe directly. In this step, we take the candidate edge objects of the thin-walled parts that were still visible before the loss of the front end of the target occlusion area and those that became visible again after the reappearance of the rear end of the target occlusion area. Using the edge positions at these two ends as constraints, we construct a candidate completion trajectory connecting the two ends, used to estimate the position evolution of the candidate edge objects of the thin-walled parts within the target occlusion area. The candidate completion trajectory refers to the trajectory to be verified obtained by estimating the position of the occluded candidate edge objects of the thin-walled parts within the target occlusion area; it has not yet been accepted. The specific construction method of the candidate completion trajectory will be elaborated in the corresponding subsequent sections.

[0025] The candidate completion trajectory is back-end verified based on the candidate edge objects of the thin-walled part that are re-emerged. The back-end verification results, the target occlusion area and the candidate edge objects of the thin-walled part are then aggregated to generate anomaly recognition results.

[0026] It should be noted that the candidate completion trajectory is estimated only based on the edge positions at both ends of the target occlusion interval. Whether it smoothly connects with the actual re-emerged thin-walled part candidate edge object at the rear end of the target occlusion interval still needs to be verified. Therefore, this step uses the actual position evolution of the thin-walled part candidate edge object corresponding to the re-emergence process as a reference to check the consistency of direction and displacement of the candidate completion trajectory near the rear end of the target occlusion interval. Only when the check passes is the candidate completion trajectory accepted as the back-end verification result. Subsequently, the back-end verification result, the target occlusion interval, and the corresponding thin-walled part candidate edge object are aggregated to generate an anomaly identification result carrying anomaly time information and object attribution information. The specific criteria for back-end verification and the specific aggregation method of anomaly identification results will be elaborated in the corresponding sections later.

[0027] This feature addresses the problem in the prior art where the trajectory obtained after subsequent completion falls on the wrong object or wrong interval, ultimately causing deviations in the start and end positions of the abnormal vibration time window and the object attribution. This is because existing solutions directly use the completed trajectory for abnormal segment extraction after obtaining it, lacking a step of checking the completion result against the real re-emergence edges. Once there is a deviation in the establishment of the target occlusion interval or the construction of the candidate completed trajectory, the deviation will be passed on to the anomaly identification result without correction. This embodiment uses the real edges of the re-emergence process to check the back end of the candidate completed trajectory, ensuring that the completed result must be consistent with the candidate edge object of the thin-walled part that is actually reproduced after the occlusion ends in terms of direction and displacement before it is accepted. This intercepts the completion deviation before anomaly identification, which is something that cannot be achieved by increasing the order of completion fitting in existing solutions that directly accept the completed trajectory.

[0028] Unlike existing technologies that use the disappearance and reappearance of target outlines in the image as occlusion events and directly proceed to trajectory compensation and abnormal segment extraction, the core difference of this embodiment lies not in the use of processing video data or CNC machining execution data as inputs, but in how to organize the relationship between these features: First, the dynamic proximity processing area and directional continuity corresponding to the tool's spatial movement process are used to jointly lock a unique candidate edge object of the thin-walled part as the analysis object. Then, the complete falling of the missing process time interval into the corresponding time sequence range of the CNC machining execution data is used as a necessary condition for the occlusion to be established, and the candidate completion trajectory is constructed by the real edge constraints at both ends of the target occlusion interval. Finally, the candidate completion trajectory is back-end verified by the real edge of the re-reappearance process before the abnormal identification result is generated. It is this organizational relationship—establishing internal anchor points based on CNC timing, extracting the two ends of the edges around the anchor points according to their sequence, and then re-examining the edges to complete the results—that allows the start and end positions of abnormal time windows and the object attribution to be constrained to the actual range caused by tool occlusion and the actual thin-walled inner sidewall edge. This effect is difficult to obtain by existing schemes that rely solely on changes in the appearance of the image edges, through adjusting parameters such as edge detection sensitivity, occlusion frame threshold, or completion fitting order, or through simple combinations.

[0029] Through the above technical solution, this embodiment uses CNC machining execution data to provide temporal support for the missing process of candidate edge objects of thin-walled parts in machining video data and then uses the re-display edge detection to complete the results. Unlike the conventional method of directly treating the disappearance and reproduction of the image outline as an occlusion event, this embodiment processes the analysis at the level of establishing the analysis object and the timing of occlusion, rather than at the level of edge detection parameters. This ensures that the target occlusion region and candidate completion trajectory generated are constrained to the real tool occlusion region and the real thin-walled inner wall edge. The target occlusion region and the candidate edge objects of thin-walled parts will be used as anchor objects and object attribution criteria in the subsequent occlusion trajectory completion and anomaly recognition result aggregation stages.

[0030] This step further elaborates on the aforementioned process of determining the dynamic proximity processing area.

[0031] Extract the tool trajectory segments containing three-dimensional coordinate changes from the CNC machining execution data, and divide the machining video data into discrete video frame sequences according to timestamps; Establish a time mapping relationship between the acquisition time of discrete video frame sequences and the execution time of tool machining trajectory segments, and project the three-dimensional coordinates of tool machining trajectory segments onto the two-dimensional pixel coordinate system of the corresponding discrete video frames according to the time mapping relationship to obtain the projected tool profile. Using the geometric center of the projected tool profile as a reference, the dynamic pixel distance is expanded outward to generate a dynamic proximity processing region.

[0032] It should be noted that, as Figure 2 As shown, the process of determining the dynamic proximity processing area is as follows: A segment showing the change in the tool's three-dimensional coordinates over time is extracted from the CNC machining execution data as the tool machining trajectory segment. This segment refers to a continuous running segment where the tool has an effective feed displacement in the machine tool coordinate system. The machining video data is divided into a sequence of discrete video frames arranged chronologically based on their frame-by-frame acquisition timestamps. A time mapping relationship is established between the acquisition time of the discrete video frames and the execution time of the tool machining trajectory segment, ensuring that any discrete video frame corresponds to the tool's three-dimensional coordinates at the same moment within the tool machining trajectory segment. Based on pre-calibrated camera projection parameters, the tool's three-dimensional coordinates at that moment are projected onto the two-dimensional pixel coordinate system of the discrete video frame, obtaining the projected tool outline in the image. Using the geometric center of the projected tool outline as a reference, a dynamic pixel distance is extended outwards in all directions of the image. The area covered by this extension is the dynamic proximity processing area corresponding to the discrete video frame. The dynamic proximity processing area moves frame by frame with the tool's projected position, thus representing a local processing range that dynamically changes within the image as the tool moves spatially.

[0033] To further explain, the time mapping relationship establishes a simultaneous correspondence between the acquisition time of discrete video frames and the execution time of the tool machining trajectory segment. When the parsing frame rate of the machining video data is inconsistent with the output timing step of the CNC machining execution data, for discrete video frames falling between two adjacent CNC machining execution data records, the tool three-dimensional coordinates of the two adjacent records are linearly interpolated according to the proportion of their acquisition time between the execution times of the two adjacent records. The tool three-dimensional coordinates at the corresponding moment of the discrete video frame are then projected, thereby avoiding the misalignment between the projected tool outline and the actual tool position due to the inconsistency of the timing steps of the two types of data.

[0034] It should be further explained that the outward expansion of the dynamic pixel distance is not a fixed value, but increases with the tool feed speed at the corresponding moment of the discrete video frame. Since a higher tool feed speed results in greater tool displacement in the frame between adjacent discrete video frames, and a larger uncertainty range for the projected tool outline between frames, the dynamic pixel distance is set to increase monotonically with the tool feed speed. This allows the dynamic proximity processing area to be appropriately enlarged at high feed speeds and appropriately tightened at low feed speeds. Its value ranges from 0.5 to 3 times the radius of the circumcircle of the projected tool outline, with a typical value of 1.5 times the radius of the circumcircle. This value ensures that the dynamic proximity processing area can completely cover the local thin-walled inner sidewall where the tool actually operates, without being too large to include irrelevant background boundaries.

[0035] It should be noted that regarding the data boundaries of discrete video frame sequences: when there are individual discrete video frames in the processing video data that are missing acquisition timestamps or have damaged images, for such discrete video frames located in the visible process or re-display process, the position is bridged by the projected tool outline position of the adjacent visible discrete video frames before and after them. Within the bridging range, the frame-by-frame generation of the dynamic neighboring processing area is not interrupted; when the number of consecutive missing or damaged discrete video frames exceeds the tolerable range limited by the subsequent tolerance frame length, the processing corresponding to the current tool processing trajectory segment is terminated, and the time mapping relationship is re-established with new visible discrete video frames after data recovery.

[0036] Through the above technical solution, this embodiment generates a dynamic proximity processing area by adaptively expanding the area with the tool feed speed based on the geometric center of the projected tool contour. Unlike the conventional method of uniformly processing the edges of the entire image or a fixed rectangular region of interest, this embodiment processes the area at the level of positioning the processing range based on the actual three-dimensional coordinates of the tool. This limits the subsequent edge extraction to the local image where the tool is actually operating. The generated dynamic proximity processing area will be used as the spatial range for performing directional gradient convolution in the directional continuity edge extraction step.

[0037] This step further expands the process of extracting candidate edge objects of thin-walled parts with directional continuity within the dynamic proximity processing area, and its execution scope is limited to the aforementioned dynamically proximity processing area.

[0038] Extract the current tool feed rate and spindle speed from the CNC machining execution data, and perform parameter weighted summation on the tool feed rate and spindle speed to generate a dynamic deflection threshold; The inter-frame search radius is calculated based on the scalar value of the tool feed rate; Perform directional gradient convolution within the dynamic proximity processing region of the current discrete video frame to generate local edges, and extract the pixel gradient distribution of the local edges to generate the tangent angle of the current frame. Centered on local edges, the matching space is delineated in adjacent discrete video frames based on the inter-frame search radius, and the local edges of adjacent frames with topological homology are locked in the matching space, thereby extracting the corresponding tangent angles of adjacent frames. Calculate the difference between the tangent angle of the current frame and the tangent angle of the adjacent frame, and output the local edge as a candidate edge object of thin-walled part with directional continuity when the difference is less than the dynamic deflection threshold.

[0039] It should be noted that, as Figure 3 As shown, the process of extracting candidate edge objects for thin-walled parts with directional continuity involves independent formulas and multiple steps with independent inputs and outputs, which are expanded into sub-steps as follows: The tool feed rate and spindle speed at the corresponding moment of the current discrete video frame are read from the CNC machining execution data. The tool feed rate and spindle speed are summed by parameter weighting to generate a dynamic deflection threshold, which serves as the allowable deviation for subsequent determination of whether the direction is continuous between adjacent frames. The input is the tool feed rate and spindle speed, and the output is the dynamic deflection threshold.

[0040] For example, the expression for calculating the dynamic deflection threshold is: ; in, For dynamic deflection threshold, For tool feed rate, Main spindle speed This is the weighting coefficient corresponding to the tool feed rate. The weighting coefficients are for the spindle speeds. All the above data have been normalized during the calculation.

[0041] Additionally, it should be noted that because the higher the tool feed rate and spindle speed, the greater the change in the true direction of candidate edge objects for thin-walled parts between adjacent discrete video frames, the dynamic deflection threshold is set to increase with weighted average of tool feed rate and spindle speed, with a value ranging from 3 to 30 degrees, and a typical value of 10 degrees; the weighting coefficient... and Based on on-site calibration data, the dynamic deflection threshold is designed to cover the fluctuations in the true direction within the commonly used feed and speed range, preventing non-homogeneous edges from being detected. Since the tool feed rate and spindle speed have different dimensions, they are converted to the same angular dimension through their respective weighting coefficients before being summed, thereby ensuring the dynamic deflection threshold is maintained. The dimensions are consistent with the difference in the angle of the subsequent tangent.

[0042] The inter-frame search radius is calculated based on the scalar value of the tool feed rate. The inter-frame search radius refers to the pixel range searched outward from the local edge when defining the matching space in adjacent discrete video frames; the input is the tool feed rate, and the output is the inter-frame search radius. Since the larger the tool feed rate, the greater the pixel displacement of the candidate edge objects of thin-walled parts between adjacent discrete video frames, the inter-frame search radius increases with the scalar value of the tool feed rate, and its value ranges from 5 pixels to 60 pixels, with a typical value of 20 pixels.

[0043] Perform directional gradient convolution within the dynamic proximity processing region of the current discrete video frame to generate local edges, and extract the pixel gradient distribution of the local edges to generate the tangent angle of the current frame. The local edge refers to the edge line segment connected by gradient responses within the dynamic proximity processing region, and the tangent angle of the current frame refers to the orientation angle of the local edge in the current discrete video frame. The input is the dynamic proximity processing region of the current discrete video frame, and the output is the local edge and its tangent angle of the current frame.

[0044] Using the local edge obtained in the previous text as the center, the matching space is delineated in adjacent discrete video frames based on the inter-frame search radius. Within the matching space, the local edges of adjacent frames that are topologically homologous to the current local edge are locked. Then, the tangent angles of adjacent frames are extracted from the local edges of adjacent frames. Topological homologous means that the local edges of two frames correspond to each other in terms of direction, length and endpoint connection relationship, which can be determined as the appearance of the same thin-walled inner sidewall edge in adjacent frames. The input is the local edge, the inter-frame search radius and the adjacent discrete video frames. The output is the local edge of the adjacent frame and its adjacent frame tangent angle.

[0045] Calculate the difference between the tangent angle of the current frame and the tangent angle of the adjacent frame, and compare this difference with the dynamic deflection threshold. When the difference is less than the dynamic deflection threshold, it is determined that the current local edge is directionally continuous between adjacent frames, and the local edge is output as a candidate edge object of thin-walled part with directional continuity. When the difference is not less than the dynamic deflection threshold, it is determined that the local edge does not have directional continuity, and it is discarded and not output as a candidate edge object of thin-walled part. The input is the tangent angle of the current frame, the tangent angle of the adjacent frame, and the dynamic deflection threshold. The output is a candidate edge object of thin-walled part with directional continuity or a discard conclusion.

[0046] It should be noted that the spatial range of the directional gradient convolution performed in this step is the dynamic proximity processing region generated by the spatiotemporal alignment and dynamic proximity processing region determination step. Since this spatial range has been constrained to the local area of ​​the thin-walled inner sidewall where the tool actually operates, the directional continuity determination in this step does not need to distinguish between the background boundary and the target edge in the entire screen range, thereby reducing the possibility that non-target contours are mistakenly selected as candidate edge objects of thin-walled parts.

[0047] The core concept of this embodiment for extracting candidate edge objects of thin-walled parts with directional continuity lies in using a dynamic deflection threshold obtained by weighting the tool feed rate and spindle speed as the threshold for determining directional continuity between adjacent frames. It also uses the inter-frame search radius calculated from the tool feed rate to lock topologically homogeneous edges in adjacent frames, thereby constraining the analysis object to a single, continuous, and adaptive thin-walled inner wall edge that adapts to machining parameters. The difference from existing technologies is that existing methods using full-screen edge detection plus fixed threshold matching do not change with machining parameters, easily missing true edges at high feed rates or high speeds and mistakenly including interfering edges at low speeds. This embodiment uses machining parameters to drive the directional continuity determination threshold and search range, which can replace existing fixed threshold edge matching and full-screen region of interest edge tracking technologies.

[0048] Through the above technical solution, this embodiment generates a dynamic deflection threshold by weighting the processing parameters and locks the directionally continuous local edges between adjacent frames by using topological homology constraints. Unlike the conventional method of performing edge matching across the entire screen with a fixed threshold, this embodiment processes the directional continuity determination threshold by making it adaptive with the tool feed rate and spindle speed. The output thin-walled part candidate edge objects with directional continuity will be used as the only analysis object for tracking the existing state frame by frame in the occlusion interval determination process.

[0049] This step further elaborates on the process of sequentially presenting the visible, missing, and re-emerging processes of candidate edge objects for thin-walled parts.

[0050] The frame rate of the processed video data is divided by the timing step of the CNC machining execution data to generate the tolerance frame length. Extract the initial pixel structure features when the candidate edge object of the thin-walled part is first tracked, and normalize the features to generate a dynamic matching baseline; In a discrete video frame sequence, the structural similarity between the first candidate edge object of the thin-walled part and the current candidate edge object of the thin-walled part is calculated, and a temporal sliding window is constructed according to the tolerance frame length before performing a frame-by-frame search: When the structural similarity of each frame within the time-domain sliding window is greater than the dynamic matching baseline, the time-domain sliding window is marked as a visible process. When the structural similarity of each frame within the temporal sliding window is no greater than the dynamic matching baseline after the visible process, the temporal sliding window is marked as a missing process. When the structural similarity of each frame within the temporal sliding window is greater than the dynamic matching baseline again after the missing process, the temporal sliding window is marked as a re-display process.

[0051] It should be noted that the process of identifying visible, missing, and re-emerging processes is as follows: The reciprocal of the temporal step size of the CNC machining execution data is taken to obtain the corresponding unit time step. The resolution frame rate of the machining video data is divided by this unit time step to obtain the tolerance frame length. The tolerance frame length refers to the number of discrete video frames corresponding to one temporal step size of the CNC machining execution data, numerically equal to the product of the resolution frame rate and the temporal step size, used to determine the length of the temporal sliding window. The initial pixel structure features of the thin-walled part candidate edge object in the image when it is first tracked are extracted and normalized to generate a dynamic matching baseline. The dynamic matching baseline is a normalized structural similarity reference used to determine whether the current thin-walled part candidate edge object can still be considered visible. The structural similarity between the first thin-walled part candidate edge object and the current thin-walled part candidate edge object is calculated frame by frame in the discrete video frame sequence. A temporal sliding window is constructed with the tolerance frame length as its length, and the search is performed by sliding the temporal sliding window frame by frame along the discrete video frame sequence. A time-domain sliding window is a time window whose length is equal to the tolerance frame length, which advances frame by frame in sequence and is used to make an overall judgment on the structural similarity of several consecutive frames.

[0052] When the structural similarity of each frame within the time-domain sliding window is greater than the dynamic matching baseline, the time-domain sliding window is marked as a visible process; when, after the visible process, the structural similarity of each frame within the time-domain sliding window is not greater than the dynamic matching baseline, the time-domain sliding window is marked as a missing process; when, after the missing process, the structural similarity of each frame within the time-domain sliding window is again greater than the dynamic matching baseline, the time-domain sliding window is marked as a re-display process.

[0053] To further clarify, the marking of the three temporal states is determined sequentially according to their order on the timeline: when the structural similarity of each frame within a time-domain sliding window is greater than the dynamic matching baseline, it indicates that the thin-walled part candidate edge object remains visible within that time-domain sliding window, and the time-domain sliding window is marked as a visible process; after a visible process has occurred, if the structural similarity of each frame within a time-domain sliding window is not greater than the dynamic matching baseline, it indicates that the thin-walled part candidate edge object remains invisible within that time-domain sliding window, and the time-domain sliding window is marked as a missing process; after a missing process has occurred, if the structural similarity of each frame within a time-domain sliding window is again greater than the dynamic matching baseline, it indicates that the thin-walled part candidate edge object becomes visible again, and the time-domain sliding window is marked as a re-emergence process. Using a time-domain sliding window of tolerance frame length for overall determination, rather than single-frame determination, ensures that fluctuations in structural similarity caused by instantaneous noise in individual frames are not misjudged as state transitions, thereby guaranteeing the stability of the boundaries of the three states—visible process, missing process, and re-emergence process—on the timeline.

[0054] It should be noted that the tolerance frame length is obtained by dividing the parsing frame rate of the processed video data by the temporal step size of the CNC machining execution data. This tolerance frame length also serves as the basis for limiting the number of consecutive missing frames that can be tolerated in the aforementioned discrete video frame sequence data boundary, so that the judgment scale of the temporal sliding window and the bridging scale of the discrete video frame sequence are coordinated with each other, thereby processing the tolerance of image missing and the judgment of state switching at the same time scale.

[0055] Through the above technical solution, this embodiment uses a time-domain sliding window of tolerance frame length to make an overall judgment on the structural similarity relative to the dynamic matching baseline, thereby sequentially dividing the visible process, the missing process, and the re-display process. Unlike the conventional method of directly determining the start and end of occlusion by whether a single frame edge is detected, this embodiment processes the process at the level of overall similarity of continuous frames rather than single frame detection as the basis for state switching. The time interval corresponding to the marked missing process will be used as the endorsement interval for the occlusion interval determination step to compare with the corresponding time sequence range of the CNC machining execution data.

[0056] This step further elaborates on the process of extracting candidate edge objects of thin-walled parts at the front and rear ends of the target occlusion area and constructing candidate completion trajectories.

[0057] Based on the start and end times of the target occlusion interval, extract the front-end visible frame group adjacent to the front end of the missing process and the back-end re-display frame group adjacent to the front end of the re-display process from the processed video data. Two-dimensional pixel coordinates of candidate edge objects of thin-walled parts in the front visible frame group and the back re-display frame group are extracted respectively to form front coordinate set and back coordinate set; Two sets of pixel sequences are obtained from the end of the front coordinate set and the beginning of the back coordinate set. The distance between the sequences, the difference in tangential angle, and the time span across frames are calculated. Spline fitting is performed on the distance between the sequences, the difference in tangential angle, and the time span across frames to generate a discrete coordinate evolution sequence within the target occlusion interval. The discrete coordinate evolution sequence is then used as a candidate completion trajectory for output.

[0058] It should be noted that the process of constructing the candidate completion trajectory is as follows: Based on the start and end times of the target occlusion interval, several discrete video frames located before the missing process and adjacent to the front end of the missing process are extracted from the processed video data as the front visible frame group, and several discrete video frames located within the re-display process and adjacent to the front end of the re-display process are extracted as the back re-display frame group; the two-dimensional pixel coordinates of the thin-walled part candidate edge objects in the front visible frame group and the back re-display frame group are extracted respectively, the former is collected into the front coordinate set, and the latter is collected into the back coordinate set; the pixel sequence of the front coordinate set near the end of the target occlusion interval is taken from the back coordinate set near the target occlusion interval. For the pixel sequence at the head of one side of the occluded area, calculate the inter-sequence distance, tangential angle difference, and frame span between these two sets of pixel sequences. The inter-sequence distance refers to the positional interval between the two sets of pixel sequences in the image. The tangential angle difference refers to the angle difference between the directions of the two sets of pixel sequences. The frame span refers to the length of time spanned by the target occlusion area. Using the inter-sequence distance, tangential angle difference, and frame span as constraints, perform spline fitting on the position evolution of the candidate edge objects of the thin-walled part within the target occlusion area to obtain a discrete coordinate evolution sequence arranged in time within the target occlusion area. This discrete coordinate evolution sequence is output as the candidate completion trajectory. The discrete coordinate evolution sequence refers to the ordered set of two-dimensional pixel coordinates of the candidate edge objects of the thin-walled part estimated at each time step within the target occlusion area.

[0059] To further explain, spline fitting uses the position and orientation of the last pixel sequence in the front coordinate set as the constraint at the beginning of the target occlusion interval, and the position and orientation of the first pixel sequence in the back coordinate set as the constraint at the end of the target occlusion interval. This ensures that the generated discrete coordinate evolution sequence connects with the positions and orientations of the thin-walled part candidate edge objects in the front visible frame group and the back re-display frame group at both ends of the target occlusion interval. The distance between sequences and the time span across frames jointly determine the position transition rate of the discrete coordinate evolution sequence within the target occlusion interval, while the tangential angle difference determines its orientation transition mode. This results in the candidate completion trajectory being continuously connected at both ends rather than simply connected by a straight line.

[0060] It should be noted that the applicable prerequisites for candidate completion trajectory construction are as follows: the front-end visible frame group and the back-end re-display frame group on which spline fitting is based must correspond to the same thin-walled inner sidewall edge, that is, the thin-walled part candidate edge objects at both ends must maintain a continuous correspondence in the sense of topological homology; when the thin-walled part candidate edge object re-displaying in the back-end re-display frame group does not satisfy the topological homology with the front-end visible frame group, indicating that the corresponding edge before and after occlusion is not the same, the candidate completion trajectory construction for the current target occlusion interval is terminated, and the target occlusion interval is re-submitted to the occlusion interval determination step to re-establish the correspondence between the visible process, the missing process and the re-display process according to the new thin-walled part candidate edge objects.

[0061] Through the above technical solution, this embodiment uses the position and orientation of the candidate edge objects of thin-walled parts that are actually visible at both ends of the target occlusion interval as constraints to perform spline fitting on the position evolution within the interval. Unlike the conventional method of directly connecting the two ends of the occlusion interval with a straight line or low-order polynomial, this embodiment processes the transition within the interval by using the position and orientation of the actual edges at both ends as constraints. The output candidate completion trajectory will be used as the object to be inspected in the anomaly identification result output stage to accept the real edge back-end verification of the re-display process.

[0062] This step further elaborates on the backend verification of the candidate completion trajectory based on the candidate edge objects of the thin-walled part that reappear.

[0063] Video frames whose acquisition time belongs to the re-display process are extracted from the processed video data to form a re-display frame sequence; Extract the edge pixel coordinates of candidate edge objects of thin-walled parts in the re-display frame sequence and concatenate them according to the acquisition time to generate the re-display edge path; Extract the rear trajectory segment that connects to the rear end of the target occlusion region from the candidate completed trajectory, and extract the starting path segment with the same frame number as the rear trajectory segment from the re-display edge path; Calculate the difference in tangent direction and the difference in displacement per unit time between the rear trajectory segment and the starting path segment; The re-display frame sequence and the dynamic proximity processing region are synthesized to generate a verification boundary. When both the tangent direction difference and the unit time displacement difference fall into the verification boundary, the candidate completed trajectory is determined to pass the back-end verification. Then, the candidate completed trajectory that passes the back-end verification is output as the back-end verification result.

[0064] It should be noted that, as Figure 4 As shown, the backend verification process includes multiple steps with independent inputs and outputs, which are broken down into sub-steps as follows: Video frames whose acquisition time belongs to the re-display process are extracted from the processed video data to form a re-display frame sequence; the input is the time interval between the processed video data and the re-display process, and the output is the re-display frame sequence.

[0065] Extract the edge pixel coordinates of the candidate edge objects of the thin-walled part in the re-display frame sequence, and concatenate them according to the acquisition time to generate the re-display edge path. The re-display edge path refers to the frame-by-frame position trajectory of the candidate edge objects of the thin-walled part after the occlusion ends. The input is the re-display frame sequence, and the output is the re-display edge path.

[0066] The candidate completed trajectory is truncated to the rear end of the target occlusion region as the rear trajectory segment, and the same frame number as the rear trajectory segment is truncated from the re-display edge path as the starting path segment, so that the two segments correspond to each other in terms of frame number and time position; the input is the candidate completed trajectory and the re-display edge path, and the output is the rear trajectory segment and the starting path segment.

[0067] Calculate the difference in tangent direction and the difference in displacement per unit time between the rear trajectory segment and the starting path segment. The difference in tangent direction refers to the angle difference between the directions of the two trajectory segments, and the difference in displacement per unit time refers to the difference between the position changes of the two trajectory segments per unit time. The input is the rear trajectory segment and the starting path segment, and the output is the difference in tangent direction and the difference in displacement per unit time.

[0068] The re-display frame sequence and the dynamic proximity processing area are synthesized to generate a verification boundary. The verification boundary refers to the allowable range of tangent direction difference and unit time displacement difference, which is jointly determined by the actual fluctuation range of the thin-walled part candidate edge object during the re-display process and the dynamic proximity processing area. When both tangent direction difference and unit time displacement difference fall within the verification boundary, the candidate completion trajectory is determined to have passed the back-end verification, and the candidate completion trajectory that has passed the back-end verification is output as the back-end verification result. When either the tangent direction difference or the unit time displacement difference falls outside the verification boundary, the candidate completion trajectory is determined to have failed the back-end verification, and it is not output as the back-end verification result. The corresponding target occlusion interval is marked as the interval to be verified. The inputs are tangent direction difference, unit time displacement difference and verification boundary, and the output is the back-end verification result or the mark to be verified.

[0069] Additionally, regarding the handling of cases that fail the backend verification: when a candidate completion trajectory fails the backend verification, it indicates that the part of the candidate completion trajectory near the back end of the target occlusion area is inconsistent with the thin-walled part candidate edge object that is actually reproduced during the re-display process in terms of direction or displacement. In this case, the candidate completion trajectory will not be included in the subsequent anomaly identification result aggregation, and the target occlusion area marked as the area to be verified will no longer generate anomaly time segments based on it, thereby avoiding the contamination of the anomaly identification results by completion results with inconsistent direction or displacement.

[0070] It should be noted that the dynamic proximity processing region used to generate the verification boundary in this step comes from the spatiotemporal alignment and dynamic proximity processing region determination step; the candidate completion trajectory used to extract the back-end trajectory segment in this step comes from the candidate completion trajectory construction step; and the re-emergence process used to generate the re-emergence edge path in this step comes from the three-stage temporal state recognition step. The synthesis and correspondence of these three aspects make the back-end verification subject to the joint constraints of the actual tool action space range, the completion result, and the real re-emergence edge. This results in the additional beneficial effect that the candidate completion trajectory is only accepted when it matches the real situation in terms of space, direction, and displacement.

[0071] The core concept of this embodiment regarding the back-end verification of candidate completed trajectories lies in using the re-emergence edge path that is actually reproduced during the re-emergence process as a reference and the verification boundary synthesized by the re-emergence frame sequence and the dynamic proximity processing region as the allowable range to perform a double consistency check on the tangent direction difference and unit time displacement difference of the back-end of the candidate completed trajectory. The difference from the prior art is that the existing scheme directly uses the completed trajectory for abnormal segment extraction without checking the completion result. Once there is a deviation in the occlusion area or the completed trajectory, it is passed to the result without correction. This scheme uses the edge that is actually reproduced after the occlusion ends to check the back-end of the completion result, which can replace the existing trajectory compensation technology that directly accepts the completed trajectory and only evaluates it based on the fitting residual.

[0072] Through the above technical solution, this embodiment performs a double check on the difference in tangent direction and the difference in displacement per unit time at the back end of the candidate completion trajectory by taking the re-display edge path as a reference to fall into the verification boundary. Unlike the conventional method of directly accepting the completion trajectory, this embodiment processes the results at the level of truly reproducing the edge back-check completion results after the occlusion ends. The output back-end verification results will be used as the accepted trajectory in the anomaly identification result aggregation stage to rearrange the trajectory time-series coordinate sequence.

[0073] This step further elaborates on the frequency domain decomposition process in the anomaly identification results generated by aggregating the backend verification results, target occlusion regions, and candidate edge objects of thin-walled parts.

[0074] The two-dimensional coordinates of the back-end verification results are rearranged according to the target occlusion interval to generate a trajectory time-series coordinate sequence; the trajectory time-series coordinate sequence is decomposed in the frequency domain to remove the feed dominant component obtained by converting the CNC machining execution data, and a jitter amplitude sequence is obtained.

[0075] It should be noted that the process of generating the jitter amplitude sequence is as follows: the two-dimensional coordinates of the candidate edge objects of the thin-walled part in the back-end verification results are rearranged according to the target occlusion interval in chronological order to obtain the trajectory time-series coordinate sequence. The trajectory time-series coordinate sequence refers to the ordered coordinate sequence of the position of the candidate edge objects of the thin-walled part within the target occlusion interval as a function of time. The trajectory time-series coordinate sequence is decomposed in the frequency domain, and the low-frequency displacement trend caused by the tool feed motion is identified as the feed dominant component and removed. The feed dominant component refers to the position change component caused by the tool feed motion rather than the vibration of the inner wall of the thin-walled part. The remaining position fluctuation after removing the feed dominant component reflects the vibration of the inner wall of the thin-walled part. The amplitude of the vibration is arranged in time to obtain the jitter amplitude sequence. The jitter amplitude sequence refers to the sequence of the vibration amplitude of the inner wall of the thin-walled part within the target occlusion interval as a function of time.

[0076] Based on the machining video data, the tool feed rate is converted into feed displacement per unit frame, and the spindle speed is converted into the number of rotations per unit time. Based on the feed displacement per unit frame, the low-frequency feed displacement trend is separated from the trajectory time-series coordinate sequence, and combined with the number of rotations per unit time, the rotational speed-related periodic components are identified to generate a jitter amplitude sequence.

[0077] To further clarify, the removal of the feed-dominant component is based on two quantities calculated from the CNC machining execution data: the tool feed rate is converted into unit frame feed displacement based on the parsing frame rate of the machining video data, where unit frame feed displacement refers to the positional change caused by tool feed between two adjacent discrete video frames; the spindle speed is converted into the number of rotations per unit time based on the parsing frame rate of the machining video data, where the number of rotations per unit time refers to the number of spindle rotations per unit time. Based on the unit frame feed displacement, the low-frequency feed displacement trend that is consistent with the tool feed direction and whose rate of change matches the unit frame feed displacement is separated from the trajectory time-series coordinate sequence and removed as the feed-dominant component. Based on the number of rotations per unit time, the periodic component associated with the spindle rotation cycle is identified, so that the jitter amplitude sequence after removing the feed-dominant component does not contain the trend displacement caused by tool feed, but retains the real vibration periodic component associated with the speed.

[0078] It should be noted that the tool feed rate and spindle speed used in this step to calculate the unit frame feed displacement and number of rotations per unit time, and the tool feed rate and spindle speed used in the direction continuity edge extraction step to generate the dynamic deflection threshold and inter-frame search radius, are taken from the same CNC machining execution data. This ensures that the same set of machining parameters remains consistent in the edge extraction and vibration separation steps, thereby producing the additional beneficial effect of the removal of the feed-dominant component and the determination of edge direction continuity being coordinated based on the same machining state.

[0079] Through the above technical solution, this embodiment separates the low-frequency feed displacement trend by converting the unit frame feed displacement from the CNC machining execution data and identifies the rotational speed-related periodic component by combining the number of rotations per unit time. Unlike the conventional method of treating the overall fluctuation of the completed trajectory as vibration indiscriminately, this embodiment processes the displacement caused by feed and the displacement caused by vibration in the frequency domain. The generated jitter amplitude sequence is used as the vibration quantity sequence for calculating the root mean square amplitude and the dynamic jitter discrimination boundary in the anomaly identification result aggregation stage.

[0080] This step further elaborates on the process of aggregating the jitter amplitude sequence, the target occlusion region, and the corresponding thin-walled part candidate edge objects to generate anomaly recognition results.

[0081] The single-rotation cycle is calculated based on the current spindle speed, and a sliding evaluation window covering the single-rotation cycle is generated by combining the processing video data. The root mean square amplitude of the jitter amplitude sequence within the sliding evaluation window is calculated, and the median amplitude and interquartile range of the trajectory of all jitter amplitudes within the target occlusion interval are extracted. Then, the root mean square amplitude, median amplitude, interquartile range, unit frame feed displacement, and number of rotations per unit time are aggregated to generate a dynamic jitter discrimination boundary.

[0082] It should be noted that the process of generating dynamic jitter discrimination thresholds and determining abnormal time segments involves independent formulas and multiple steps with independent inputs and outputs, which are broken down into sub-steps as follows: The single-revolution cycle is calculated based on the spindle speed, which is the time required for the spindle to rotate one revolution. A sliding evaluation window with a length covering the single-revolution cycle is generated by combining the parsed frame rate of the processing video data, so that each sliding evaluation window includes at least one revolution of the spindle in time. The input is the spindle speed and the processing video data, and the output is the sliding evaluation window.

[0083] Calculate the root mean square amplitude of the jitter amplitude sequence within each sliding evaluation window. The root mean square amplitude refers to the root mean square value of the jitter amplitude within the sliding evaluation window, which is used to characterize the overall intensity of the vibration of the inner wall of the thin-walled structure within that sliding evaluation window. The input is the jitter amplitude sequence and the sliding evaluation window, and the output is the root mean square amplitude of each sliding evaluation window.

[0084] Extract the median amplitude and interquartile range of all jitter amplitudes within the target occlusion interval. The median amplitude refers to the median of all jitter amplitudes within the target occlusion interval, and the interquartile range refers to the difference between the upper and lower quartiles of all jitter amplitudes within the target occlusion interval. Aggregate the root mean square amplitude, median amplitude, interquartile range, unit frame feed displacement, and number of rotations per unit time to generate a dynamic jitter discrimination threshold. The inputs are the root mean square amplitude, median amplitude, interquartile range, unit frame feed displacement, and number of rotations per unit time, and the output is the dynamic jitter discrimination threshold.

[0085] For example, the expression for calculating the dynamic jitter discrimination threshold is: ; in, To determine the threshold for dynamic jitter, The median amplitude of the trajectory. The interquartile range of the amplitude. The unit frame feed displacement The number of rotations per unit time. The dispersion coefficients corresponding to the interquartile range of the amplitude. The coefficient corresponding to the unit frame feed displacement. The coefficient corresponds to the number of rotations per unit time. All the above data have been normalized during the calculation.

[0086] It should be further explained that, since the median amplitude and interquartile range of the amplitude jointly characterize the center level and dispersion of the jitter amplitude within the target occlusion interval, while the unit frame feed displacement and the number of rotations per unit time characterize the influence of the current processing state on the normal vibration level, the dynamic jitter discrimination limit is set as a margin superimposed on the median amplitude of the trajectory, determined by the interquartile range of the amplitude and the processing state, so that it adapts to the vibration distribution and processing state within the target occlusion interval; the dispersion coefficient... The value ranges from 1.0 to 3.0, with a typical value of 1.5. (The coefficient is missing from the original text.) and Based on the on-site calibration data, the dynamic jitter discrimination limit is aligned with the upper edge of the normal vibration level under common processing conditions; the median amplitude, interquartile range, and root mean square amplitude of the trajectory are all measured in pixels, and the unit frame feed displacement is calculated using a coefficient. Number of rotations per unit time, with coefficient Converting to pixel dimensions and then adding them together ensures the dynamic jitter detection threshold. Its dimensions are consistent with those of the root mean square amplitude.

[0087] When the root mean square amplitude of adjacent sliding evaluation windows is higher than the dynamic jitter discrimination limit, and the adjacent sliding evaluation windows are connected continuously on the time axis, the start time of the first out-of-limit sliding evaluation window to the end time of the last out-of-limit sliding evaluation window is determined as an abnormal time segment; the abnormal time segment, the corresponding thin-walled part candidate edge object, the target occlusion interval and the back-end verification result are aggregated into an anomaly identification result.

[0088] The root mean square amplitude of each sliding evaluation window is compared with the dynamic jitter discrimination threshold along the time axis. When the root mean square amplitude of several adjacent sliding evaluation windows is higher than the dynamic jitter discrimination threshold and these sliding evaluation windows are connected continuously on the time axis, the time interval spanned from the start time of the first exceeding sliding evaluation window to the end time of the last exceeding sliding evaluation window is determined as an abnormal time segment. When there are no adjacent and continuously connected exceeding sliding evaluation windows, that is, when a single sliding evaluation window exceeds the limit but its adjacent sliding evaluation windows do not exceed the limit, an abnormal time segment is not determined based on the isolated exceeding sliding evaluation window, thereby avoiding misjudging a single window exceeding the limit due to occasional noise as abnormal. The input is the root mean square amplitude of each sliding evaluation window and the dynamic jitter discrimination threshold, and the output is an abnormal time segment or a conclusion of no abnormal time segment.

[0089] The abnormal time segment, the candidate edge object of the thin-walled part corresponding to the abnormal time segment, the target occlusion region, and the back-end verification result are aggregated into the anomaly recognition result; the input is the abnormal time segment, the candidate edge object of the thin-walled part, the target occlusion region, and the back-end verification result, and the output is the anomaly recognition result.

[0090] It should be noted that the jitter amplitude sequence used to determine the abnormal time segment in this step comes from the frequency domain decomposition of the jitter amplitude sequence, the back-end verification result used to aggregate comes from the candidate completion trajectory back-end verification step, and the target occlusion interval used to limit the time attribution comes from the occlusion interval determination step. The aggregation of these three results enables the anomaly identification result to simultaneously carry the start and end times of the anomaly, the true intensity of the vibration, and the object attribution information, thereby producing the additional beneficial effect of binding the start and end positions of the abnormal time window with the object attribution without misalignment.

[0091] It should be noted that the anomaly identification results output in this step are provided with specific fields in two scenarios depending on whether an abnormal time segment exists. In the case of an abnormal time segment, the anomaly identification result includes the start and end times of the abnormal time segment, the identifier of the candidate edge object of the thin-walled part corresponding to the abnormal time segment and its frame-by-frame two-dimensional pixel coordinates, the start and end times of the target occlusion interval, and the discrete coordinate evolution sequence in the back-end verification result. For example, a specific anomaly identification result is as follows: the abnormal time segment is from 1.20 seconds to 1.86 seconds, the corresponding candidate edge object identifier of the thin-walled part is the left edge of the inner wall, the target occlusion interval is from 1.00 seconds to 2.00 seconds, the back-end verification result is the frame-by-frame two-dimensional pixel coordinate evolution sequence within the target occlusion interval, and an appended root mean square magnitude sequence within the abnormal time segment. In the absence of abnormal time segments, the anomaly identification result includes the start and end times of the target occlusion interval, the identification of the corresponding thin-walled part candidate edge object, and the judgment conclusion that there are no abnormal time segments. This is used to indicate that although the target occlusion interval is confirmed to be tool occlusion, the vibration of the inner wall of the thin-walled part does not exceed the dynamic jitter discrimination limit.

[0092] The core concept of the dynamic jitter discrimination boundary and abnormal time segment determination scheme in this embodiment lies in constructing a dynamic jitter discrimination boundary that is adaptive to the distribution and working conditions by combining the median amplitude and interquartile range of the jitter amplitude trajectory within the target occlusion interval with processing state quantities. Abnormal time segments are defined by the first and last moments of multiple adjacent and continuously connected over-limit sliding evaluation windows. The difference from the prior art is that the existing schemes mostly use fixed amplitude thresholds or single-frame over-limits to directly determine anomalies, which neither adjusts with working conditions and vibration distribution, and is prone to misjudging occasional noise as anomalies. This scheme uses an adaptive discrimination boundary combined with adjacent continuous over-limit constraints to define abnormal time segments, which can replace existing fixed threshold vibration determination, single-frame anomaly triggering and other technologies.

[0093] For example, the following provides a complete numerical implementation process from establishing the target occlusion region to outputting the anomaly recognition result. Assume the parsing frame rate of the processing video data is 100 frames per second, and the timing step of the CNC machining execution data is 0.1 seconds. Then, the number of time steps per unit time corresponding to the timing step of 0.1 seconds is its reciprocal, i.e., 10 steps per second. Dividing the parsing frame rate of 100 frames per second by 10 steps per second yields a tolerance frame length of 10 frames. This result is numerically equal to the product of the parsing frame rate of 100 frames per second and the timing step of 0.1 seconds, meaning that one timing step contains exactly 10 frames. Assume the tool feed rate is 800 mm per minute and the spindle speed is 3000 rpm. Substituting these values ​​into the dynamic deflection threshold expression and taking the weighting coefficients... 0.005 degrees per millimeter per minute The dynamic deflection threshold is 0.002 degrees per revolution per minute. 0.005 multiplied by 800 plus 0.002 multiplied by 3000 equals 4.0 plus 6.0, which equals 10.0 degrees, consistent with the typical value and falling within the range of 3 to 30 degrees; the inter-frame search radius is calculated based on the tool feed rate and taken as 20 pixels, consistent with the typical value and falling within the range of 5 to 60 pixels.

[0094] In frame-by-frame tracking, the structural similarity of candidate edge objects for thin-walled parts is greater than the dynamic matching baseline in each frame between 0.00 and 0.99 seconds, and the corresponding temporal windows are marked as visible processes. From 1.00 to 1.99 seconds, the structural similarity is not greater than the dynamic matching baseline in each frame, and the corresponding temporal windows are marked as missing processes. From 2.00 seconds onwards, the structural similarity is again greater than the dynamic matching baseline in each frame, and the corresponding temporal windows are marked as re-enhanced processes. Reading the CNC machining execution data shows that the timing range of the tool passing through the adjacent position of the inner sidewall to be machined is from 0.95 seconds to 2.05 seconds. The missing process, from 1.00 to 1.99 seconds, falls completely within this timing range. Therefore, 1.00 to 2.00 seconds is established as the target occlusion interval.

[0095] Discrete video frames from 0.90 to 0.99 seconds before the target occlusion interval are taken as the front visible frame group, and discrete video frames from 2.00 to 2.09 seconds during the re-display process are taken as the back re-display frame group. Two-dimensional pixel coordinates of candidate edge objects of thin-walled parts are extracted to form front coordinate sets and back coordinate sets respectively. Spline fitting is performed with the distance between the sequence of pixel points at the end of the front coordinate set and the beginning of the back coordinate set, the difference in tangential angle, and the time span of 1.00 seconds across frames as constraints to generate a discrete coordinate evolution sequence frame by frame from 1.00 seconds to 2.00 seconds as candidate completion trajectory. The back-end trajectory segment that connects to the back end of the target occlusion interval in the candidate completion trajectory and the starting path segment with the same frame number in the re-display edge path are extracted. The tangent direction difference is calculated to be 2.1 degrees and the unit time displacement difference is 0.6 pixels per frame. Both fall within the verification boundary (the allowable range of tangent direction difference is 0 to 5 degrees and the allowable range of unit time displacement difference is 0 to 1.5 pixels per frame) synthesized by the re-display frame sequence and the dynamic proximity processing area. The candidate completion trajectory is determined to pass the back-end verification and is output as the back-end verification result.

[0096] The trajectory time-series coordinate sequence is obtained by rearranging the two-dimensional coordinates of the back-end verification results according to the target occlusion interval; the tool feed rate is converted into unit frame feed displacement according to the parsed frame rate. That is, 800 mm per minute first translates to a feed displacement of approximately 13.33 mm per second, then divides by the resolution frame rate of 100 frames per second to get a feed displacement of approximately 0.133 mm per frame. Based on the calibrated camera pixel equivalent of 9 pixels per millimeter, the feed displacement per unit frame is calculated. Approximately 1.2 pixels per frame. Divide the spindle speed of 3000 revolutions per minute by 60 to convert it into the number of rotations per unit time. The rotation speed is 50 revolutions per second. The low-frequency feed displacement trend is separated based on the unit frame feed displacement, and the rotational speed-related periodic component is identified based on the number of rotations per unit time. After removing the dominant feed component, a jitter amplitude sequence is obtained. The median amplitude of the trajectory of all jitter amplitudes within the target occlusion interval is extracted. 3.0 pixels, amplitude interquartile range The value is 1.2 pixels, and the dispersion coefficient is taken. The coefficient is 1.5. For 0.5 frames, Substituting 0.01 pixels per second per revolution into the dynamic jitter discrimination boundary expression, the dynamic jitter discrimination boundary is then... Adding 1.5, multiplying by 1.2, adding 0.5, multiplying by 1.2, adding 0.01, and multiplying by 50 gives 3.0, 1.8, 0.6, and 0.5, which equals 5.9 pixels.

[0097] Based on the spindle speed, the single-rotation cycle is calculated to be 0.02 seconds, generating a sliding evaluation window covering the single-rotation cycle. The root mean square (RMS) amplitude of the jitter amplitude sequence within each sliding evaluation window is calculated. Among them, the RMS amplitudes of several adjacent sliding evaluation windows between 1.20 seconds and 1.86 seconds are 6.2 pixels, 6.5 pixels, 6.4 pixels, and 6.1 pixels respectively, all higher than the dynamic jitter discrimination threshold of 5.9 pixels and continuously connected on the time axis. The RMS amplitudes of sliding evaluation windows outside this interval are all no higher than 5.9 pixels. Therefore, the period from the start time of the first out-of-limit sliding evaluation window (1.20 seconds) to the end time of the last out-of-limit sliding evaluation window (1.86 seconds) is identified as an abnormal time segment. This abnormal time segment, the corresponding left edge of the inner wall of the candidate edge object of the thin-walled part, the target occlusion interval from 1.00 seconds to 2.00 seconds, and the back-end verification results are aggregated to generate the anomaly identification result output. In this numerical implementation process, all intermediate values ​​are consistent with the value range and typical values ​​mentioned above.

[0098] Through the above technical solution, this embodiment defines abnormal time segments and aggregates them with object attribution by using a dynamic jitter discrimination boundary that adapts to vibration distribution and processing status, combined with adjacent continuous over-limit sliding evaluation windows. This differs from the conventional processing method that triggers anomalies in a single frame with a fixed amplitude threshold. This embodiment processes the anomaly judgment boundary by adapting it to the working conditions and defining time segments with continuous over-limit constraints. The output anomaly identification result serves as the final judgment conclusion of this method for processing monitoring. It is used as the basis for determining the start and end of the abnormal time window and object attribution in subsequent anomaly warning and processing parameter intervention stages.

[0099] Example 2: Please see Figure 5 A machine tool machining vibration state recognition system based on video analytics includes: The data acquisition module is used to acquire the processing video data and CNC machining execution data of the target object; The tool proximity edge filtering module is used to perform spatiotemporal alignment between machining video data and CNC machining execution data, determine the dynamic proximity processing area corresponding to the tool's spatial movement process in the video frame, and extract candidate edge objects of thin-walled parts with directional continuity within the dynamic proximity processing area. The occlusion interval determination module is used to track candidate edge objects of thin-walled parts frame by frame in the processing video data, and when the candidate edge objects of thin-walled parts are identified to show the temporal state changes of visible process, missing process and re-showing process in sequence, and when the time interval of the missing process falls completely within the temporal range corresponding to the CNC machining execution data, the time interval corresponding to the missing process is established as the target occlusion interval. The occlusion trajectory completion module is used to extract candidate edge objects of thin-walled parts at the front and rear ends of the target occlusion area and construct candidate completion trajectories connecting the two ends accordingly. The anomaly detection result output module is used to perform backend verification on the candidate completion trajectory based on the candidate edge object of the thin-walled part that is re-displayed, and to aggregate the backend verification result, the target occlusion area and the candidate edge object of the thin-walled part to generate the anomaly detection result.

[0100] This embodiment has the same technical effects as Embodiment 1.

[0101] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The data mentioned in this application have undergone normalization and other preprocessing to unify dimensions during formula calculations.

[0102] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A machine tool machining vibration state recognition method based on video analysis, applied to the identification of abnormal time windows during the milling vibration of the inner wall of thin-walled frame parts in a vertical machining center, characterized in that... Includes the following steps: Acquire the processing video data and CNC machining execution data of the target object; The machining video data and the CNC machining execution data are spatiotemporally aligned to determine the dynamic proximity processing area corresponding to the tool spatial movement process in the video frame, and candidate edge objects of thin-walled parts with directional continuity are extracted within the dynamic proximity processing area. In the processing video data, the candidate edge objects of the thin-walled part are tracked frame by frame. When the candidate edge objects of the thin-walled part are identified to show a temporal state change of visible process, missing process and re-showing process in sequence, and when the time interval of the missing process falls completely within the temporal range corresponding to the CNC machining execution data, the time interval corresponding to the missing process is established as the target occlusion interval. Extract the candidate edge objects of the thin-walled parts at the front and rear ends of the target occlusion area and construct the candidate completion trajectory connecting the two ends accordingly; The candidate completion trajectory is back-end verified based on the candidate edge object of the thin-walled part that is re-displayed, and the back-end verification result, the target occlusion region and the candidate edge object of the thin-walled part are aggregated to generate an anomaly recognition result.

2. The machine tool processing vibration state identification method based on video analysis according to claim 1, characterized in that: The spatiotemporal alignment of the machining video data and the CNC machining execution data, and the determination of the dynamic proximity processing area corresponding to the tool spatial movement process in the video frame, specifically includes: Extract the tool machining trajectory segment containing three-dimensional coordinate changes from the CNC machining execution data, and divide the machining video data into discrete video frame sequences according to timestamps; Establish a time mapping relationship between the acquisition time of the discrete video frame sequence and the execution time of the tool machining trajectory segment, and project the three-dimensional coordinates of the tool machining trajectory segment onto the two-dimensional pixel coordinate system of the corresponding discrete video frame according to the time mapping relationship to obtain the projected tool outline; Using the geometric center of the projection tool profile as a reference, the dynamic pixel distance is expanded outward to generate a dynamic proximity processing region.

3. The machine tool processing vibration state identification method based on video analysis according to claim 2, characterized in that: Extracting candidate edge objects of thin-walled parts with directional continuity within the dynamic proximity processing area specifically includes: Extract the current tool feed rate and spindle speed from the CNC machining execution data, and perform parameter weighted summation on the tool feed rate and spindle speed to generate a dynamic deflection threshold; The inter-frame search radius is calculated based on the scalar value of the tool feed rate; Perform directional gradient convolution within the dynamic proximity processing region of the current discrete video frame to generate local edges, and extract the pixel gradient distribution of the local edges to generate the tangent angle of the current frame. Centered on the local edge, a matching space is delineated in adjacent discrete video frames based on the inter-frame search radius, and topologically homogeneous local edges of adjacent frames are locked in the matching space, thereby extracting the corresponding tangent angles of adjacent frames. Calculate the difference between the tangent angle of the current frame and the tangent angle of the adjacent frame, and when the difference is less than the dynamic deflection threshold, output the local edge as a candidate edge object of a thin-walled part with directional continuity.

4. The machine tool processing vibration state identification method based on video analysis according to claim 2, characterized in that: The temporal state changes of the candidate edge objects of the thin-walled part, which are identified as exhibiting a visible process, a missing process, and a re-emergence process, specifically include: The frame rate of the processed video data is divided by the timing step of the CNC machining execution data to generate the tolerance frame length. Extract the initial pixel structure features when the candidate edge object of the thin-walled part is first tracked, and normalize the features to generate a dynamic matching baseline; In the discrete video frame sequence, the structural similarity between the first candidate edge object of the thin-walled part and the current candidate edge object of the thin-walled part is calculated, and a temporal sliding window is constructed according to the tolerance frame length before performing a frame-by-frame search: When the structural similarity of each frame within the temporal sliding window is greater than the dynamic matching baseline, the temporal sliding window is marked as a visible process. When the structural similarity of each frame within the temporal sliding window is no greater than the dynamic matching baseline after the visible process, the temporal sliding window is marked as a missing process. When the structural similarity of each frame within the temporal sliding window is greater than the dynamic matching baseline again after the missing process, the temporal sliding window is marked as a re-display process.

5. The machine tool processing vibration state identification method based on video analysis according to claim 1, characterized in that: Extracting candidate edge objects of thin-walled parts at the front and rear ends of the target occlusion region and constructing candidate completion trajectories connecting the two ends accordingly specifically includes: Based on the start and end times of the target occlusion interval, extract the front-end visible frame group adjacent to the front end of the missing process and the back-end re-display frame group adjacent to the front end of the re-display process from the processed video data. Two-dimensional pixel coordinates of the candidate edge objects of the thin-walled part in the front-end visible frame group and the back-end re-display frame group are extracted respectively to form a front-end coordinate set and a back-end coordinate set. Two sets of pixel sequences are obtained based on the end of the front-end coordinate set and the beginning of the back-end coordinate set. The sequence distance, tangential angle difference, and cross-frame time span of the two sets of pixel sequences are calculated. Spline fitting is performed on the sequence distance, tangential angle difference, and cross-frame time span to generate a discrete coordinate evolution sequence within the target occlusion interval. The discrete coordinate evolution sequence is then output as a candidate completion trajectory.

6. The machine tool processing vibration state identification method based on video analysis according to claim 1, characterized in that: The backend verification of the candidate completion trajectory based on the re-emergence of the corresponding thin-walled part candidate edge object specifically includes: Video frames whose acquisition time belongs to the re-display process are extracted from the processed video data to form a re-display frame sequence; Extract the edge pixel coordinates of candidate edge objects of thin-walled parts in the re-display frame sequence and concatenate them according to the acquisition time to generate the re-display edge path; Extract the rear trajectory segment that connects to the rear end of the target occlusion region from the candidate completed trajectory, and extract the starting path segment with the same frame number as the rear trajectory segment from the re-display edge path; Calculate the difference in tangent direction and the difference in displacement per unit time between the rear trajectory segment and the starting path segment; The re-display frame sequence and the dynamic proximity processing region are synthesized to generate a verification boundary. When both the tangent direction difference and the unit time displacement difference fall into the verification boundary, the candidate completion trajectory is determined to pass the back-end verification. The candidate completion trajectory that passes the back-end verification is then output as the back-end verification result.

7. The machine tool processing vibration state identification method based on video analysis according to claim 3, characterized in that: The backend verification results, the target occlusion region, and the candidate edge objects of the thin-walled part are aggregated to generate an anomaly identification result, specifically including: The two-dimensional coordinates of the back-end verification results are rearranged according to the target occlusion interval to generate a trajectory time-series coordinate sequence; The trajectory time-series coordinate sequence is decomposed in the frequency domain to remove the feed-dominant component obtained from the CNC machining execution data, resulting in a jitter amplitude sequence. The jitter amplitude sequence, the target occlusion interval, and the corresponding thin-walled part candidate edge objects are aggregated to generate the anomaly identification result.

8. The machine tool processing vibration state identification method based on video analysis according to claim 7, characterized in that: The trajectory time-series coordinate sequence is decomposed in the frequency domain to remove the feed-dominant component obtained from the CNC machining execution data, resulting in a jitter amplitude sequence that specifically includes: The tool feed rate is converted into feed displacement per unit frame based on the machining video data, and the spindle speed is converted into the number of rotations per unit time based on the machining video data. Based on the unit frame feed displacement, the low-frequency feed displacement trend is separated from the trajectory time-series coordinate sequence, and the rotational speed-related periodic component is identified by combining the number of rotations per unit time, thereby generating a jitter amplitude sequence.

9. The machine tool processing vibration state identification method based on video analysis according to claim 8, characterized in that: The aggregation of the jitter amplitude sequence, the target occlusion region, and the corresponding thin-walled part candidate edge objects to generate the anomaly identification result specifically includes: The single-revolution cycle is calculated based on the current spindle speed, and a sliding evaluation window covering the single-revolution cycle is generated by combining the machining video data. The root mean square amplitude of the jitter amplitude sequence within the sliding evaluation window is calculated, and the median amplitude and interquartile range of the trajectory of all jitter amplitudes within the target occlusion interval are extracted. Then, the root mean square amplitude, the median amplitude, the interquartile range, the unit frame feed displacement, and the number of rotations per unit time are aggregated to generate a dynamic jitter discrimination threshold. When the root mean square amplitude of adjacent sliding evaluation windows is higher than the dynamic jitter discrimination limit, and the adjacent sliding evaluation windows are connected continuously on the time axis, the start time of the first over-limit sliding evaluation window to the end time of the last over-limit sliding evaluation window is determined as an abnormal time segment. The abnormal time segment, the corresponding thin-walled part candidate edge object, the target occlusion interval, and the back-end verification result are aggregated into an anomaly identification result.

10. A machine tool machining vibration state recognition system based on video analysis, characterized in that, include: The data acquisition module is used to acquire the processing video data and CNC machining execution data of the target object; The tool proximity edge filtering module is used to perform spatiotemporal alignment between the machining video data and the CNC machining execution data, determine the dynamic proximity processing area corresponding to the tool spatial movement process in the video frame, and extract thin-walled part candidate edge objects with directional continuity within the dynamic proximity processing area. The occlusion interval determination module is used to track the candidate edge objects of the thin-walled part frame by frame in the processing video data, and when the candidate edge objects of the thin-walled part are identified to show the temporal state changes of the visible process, the missing process and the re-display process in sequence, and when the time interval of the missing process falls completely within the temporal range corresponding to the CNC machining execution data, the time interval corresponding to the missing process is established as the target occlusion interval. The occlusion trajectory completion module is used to extract the candidate edge objects of the thin-walled parts at the front and rear ends of the target occlusion area and construct the candidate completion trajectory connecting the two ends accordingly. The anomaly detection result output module is used to perform backend verification on the candidate completion trajectory based on the candidate edge object of the thin-walled part that is re-displayed, and to aggregate the backend verification result, the target occlusion region and the candidate edge object of the thin-walled part to generate anomaly detection result.

Citation Information

Patent Citations

  • Industrial safety cross-domain collaborative analysis system based on time series data and visual data

    CN121327446A

  • Time-sequential image characteristics extraction method and device therefor and record medium for recording the same method

    JP2000011182A