Kitchen raw and cooked food material cross contamination early warning system and method based on multi-modal fusion

By using multimodal fusion technology, visible light video, thermal infrared video and depth/binocular data are spatiotemporally registered and relational pattern recognized to generate triplet event streams, solving the problem of identifying and obtaining evidence of cross-contamination transmission chains of raw and cooked food in the kitchen, and achieving reliable real-time early warning and traceability.

CN121963052APending Publication Date: 2026-05-01JIANGSU FOOD & PHARMA SCI COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU FOOD & PHARMA SCI COLLEGE
Filing Date
2026-01-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately identify the transmission chain of cross-contamination between raw and cooked food in the kitchen, resulting in insufficient interpretability of early warning results and incomplete evidence collection chains, making it difficult to support accurate accountability and regulatory record-keeping.

Method used

By using multimodal fusion technology, visible light video, thermal infrared video and depth/binocular data are spatiotemporally registered to generate a multimodal object candidate set. The triplet event stream and object identity trajectory table are output through open vocabulary video relationship pattern recognition rules to form a cross-contamination evidence chain and perform consistency verification to trigger real-time early warning.

Benefits of technology

It enables reliable and traceable early warning of cross-contamination between raw and cooked food in the kitchen, accurately identifies the path of contamination transmission and generates interpretable risk profiles, thereby improving the reliability and traceability of early warning of cross-contamination in the kitchen.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963052A_ABST
    Figure CN121963052A_ABST
Patent Text Reader

Abstract

The invention discloses a kitchen raw and cooked food material cross contamination early warning system and method based on multi-modal fusion, and relates to the technical field of computer vision, and the method comprises the steps: obtaining a visible light video, a thermal infrared video and depth / binocular data in a kitchen operation region, completing the time-space registration, determining a target candidate of a kitchen operation entity, and carrying out the early warning of the cross contamination of the kitchen raw and cooked food materials; generating a multi-modal object candidate set; and performing relation judgment on the multi-mode object candidate set according to an open vocabulary video relation mode recognition rule, and outputting a triple event stream and a corresponding object identity track table. According to the method, an open vocabulary video relation mode is adopted to recognize and output a triple event stream and an object identity track table, so that interaction relations such as contact, bearing, transmission, taking and placement stably fall to the ground in an event segment form, and therefore, sequence division and pollution candidate link retrieval in a subsequent stage are supported.
Need to check novelty before this filing date? Find Prior Art

Description

A Multimodal Fusion-Based Early Warning System and Method for Cross-Contamination of Raw and Cooked Foods in Kitchens Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a multimodal fusion-based early warning system and method for cross-contamination of raw and cooked food ingredients in the kitchen. Background Technology

[0002] With the development of smart catering and digital food safety supervision, kitchen operation monitoring is gradually shifting from manual inspection to automatic identification solutions based on visual perception. Existing technologies typically use visible light video to detect targets and track trajectories in the work area, and combine this with area segmentation rules to continuously observe objects such as knives, cutting boards, containers, and hands. Some solutions further overlay thermal infrared or depth / binocular data to enhance the stability of target localization in occluded scenarios, and achieve cross-modal data alignment through multi-view calibration and spatiotemporal registration, thereby providing a unified coordinate and time index basis for subsequent behavior recognition.

[0003] However, cross-contamination between raw and cooked food in the kitchen is a risk coupled with "object-relationship-process". Traditional methods that focus on object category or single action identification are difficult to depict the propagation chain of interactive relationships such as "contact, carrying, passing, picking up, and placing" in time series. This can easily lead to the problem of only identifying "object existence" but failing to confirm "contamination propagation path". At the same time, existing methods mostly use multimodal evidence at the target detection level, lacking a structured expression of relationship triple events and a cross-modal consistency verification mechanism. This results in insufficient interpretability of early warning results, incomplete evidence chains, and difficulty in supporting accurate accountability and regulatory record keeping. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion, which solves the problem that the transmission chain of cross-contamination between raw and cooked food in the kitchen is difficult to accurately identify and trace.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a method for early warning of cross-contamination between raw and cooked food ingredients in a kitchen based on multimodal fusion, comprising,

[0008] Visible light video, thermal infrared video and depth / binocular data are acquired in the kitchen operation area and spatiotemporal registration is completed to determine target candidates for kitchen operation entities and generate a multimodal object candidate set.

[0009] Relationship discrimination is performed on the multimodal object candidate set according to the open vocabulary video relationship pattern recognition rules, and the triple event stream and the corresponding object identity trajectory table are output.

[0010] Based on the event flow of relation triples and the object identity trajectory table, the stage sequence of the back kitchen operation process is determined and the contamination event is assembled to form a cross-contamination evidence chain;

[0011] Consistency verification is performed on the evidence chain of cross-contamination to form a traceable evidence package, triggering real-time early warning of cross-contamination in the kitchen scenario and generating a risk profile.

[0012] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food in the kitchen, the kitchen operation area includes a raw food area, a cooked food area, a washing and disinfection area, and a serving outlet.

[0013] The kitchen operation entities include processing tools, supporting equipment, operating areas, personnel hands, and food ingredients.

[0014] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food ingredients in the kitchen according to the present invention, the specific steps for completing the spatiotemporal registration are as follows:

[0015] After aligning the visible light video, thermal infrared video and depth / binocular data with timestamps, a calibration board is used to acquire multi-view calibration images and calculate the camera intrinsic parameter calibration results of the visible light camera, thermal infrared camera and depth / binocular camera. The camera intrinsic parameter calibration results are used to complete the distortion correction of each view and generate a corrected image sequence.

[0016] Collect the corresponding feature points of the same calibration plate under the view of each camera and calculate the camera extrinsic calibration results. Use the camera extrinsic calibration results to establish the coordinate mapping relationship between visible light, thermal infrared and depth / binocular data, and generate a cross-modal mapping matrix.

[0017] Extract platform point clouds from depth / stereo data and perform plane fitting to obtain operational plane extraction results. Use the operational plane extraction results to determine the platform plane coordinate system and complete the unification of multimodal coordinates. Output multimodal aligned data with unified coordinates and unified time index.

[0018] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food in the kitchen according to the present invention, the specific steps for generating the multimodal object candidate set are as follows:

[0019] On the visible light frame, thermal infrared frame and depth / binocular frame of the multimodal aligned data, target detection and instance segmentation, target segmentation and contour extraction, and depth clustering and 3D boundary extraction are performed respectively to obtain visible light candidate set, thermal infrared candidate set and depth candidate set;

[0020] Cross-modal correlation is performed between the visible light candidate set, the thermal infrared candidate set, and the depth candidate set, and short-term tracking is performed along the time dimension to generate short object trajectories and assign trajectory labels.

[0021] The 3D center point and orientation information of the object are calculated using the 3D boundary of the depth candidate set and the table plane coordinate system, and bound to the trajectory identifier of the object's short trajectory to output a multimodal object candidate set.

[0022] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food in the kitchen according to the present invention, the specific steps for determining the relationship of the multimodal object candidate set according to the open vocabulary video relationship pattern recognition rule are as follows:

[0023] Obtain the kitchen interaction relation terminology and kitchen operation entity terminology to form a relation hint set;

[0024] Extract trajectory temporal features from each short trajectory of an object in the multimodal object candidate set and determine entity category labels. At the same time, divide the short trajectory of the object into a subject trajectory set and an object trajectory set according to the entity category labels.

[0025] Calculate the temporal overlap and spatial proximity segments of any two short trajectories of objects in the subject trajectory set and the object trajectory set, and generate a set of relational spatiotemporal segments;

[0026] Visible light appearance features, thermal infrared contour features, and three-dimensional pose change features are extracted from the spatiotemporal fragment set of relations and fused to form a relation feature sequence;

[0027] The relation feature sequence and relation cue set are subjected to cross-modal contrastive matching in a visual-language joint model to generate a relation score sequence;

[0028] Perform time-consistency decoding on the relation score sequence and determine the start and end range of the relation to form relation triplet event fragments;

[0029] The event fragments of the relation triples are associated and registered with the corresponding short trajectory identifiers of the subject and the object, generating a triple event flow and a corresponding object identity trajectory table.

[0030] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food ingredients in the kitchen according to the present invention, the specific steps for obtaining the kitchen interaction relation word list and the kitchen operation entity word group library are as follows:

[0031] Construct an entity category list based on kitchen operation entities, and generate kitchen operation entity phrases;

[0032] Construct a kitchen interaction relation lexicon based on the semantics of the kitchen operation process;

[0033] Combine the entity phrases of kitchen operations with the kitchen interaction relation vocabulary to generate a relational text sequence;

[0034] Perform text encoding on relational text sequences to generate a relational semantic prototype library;

[0035] By combining the relational semantic prototype library with relational text sequences, a relational hint set is formed.

[0036] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food in the kitchen according to the present invention, the specific steps for determining the stage sequence of the kitchen operation process and assembling contamination events are as follows:

[0037] Extract tool trajectory, hand trajectory, and food trajectory from the object identity trajectory table to obtain a set of motion fragments;

[0038] The set of action segments is divided into stages according to the operational semantics to obtain a stage sequence;

[0039] Align the phase sequence with the triplet event flow in time to form a phase relationship graph;

[0040] Cross-stage link search was performed on raw food-related events and cooked food-related events in the stage relationship diagram to obtain candidate contamination links;

[0041] The candidate contamination links are assembled according to their time sequence and spatial location to form a chain of cross-contamination evidence.

[0042] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food ingredients in the kitchen according to the present invention, the specific steps for performing consistency verification on the cross-contamination evidence chain are as follows:

[0043] Extract key nodes from the cross-contamination evidence chain to generate a key frame index;

[0044] Cross-modal consistency verification is performed on the visible light evidence, thermal infrared evidence, and depth / binocular evidence indicated by the keyframe index to generate consistency markers;

[0045] The consistency markers, object identity trajectory tables, triplet event flows, and phase sequences are correlated and summarized to generate a set of evidence entries.

[0046] A link description is generated based on the chronological order of key nodes in the cross-contamination evidence chain and the connection relationship between adjacent nodes. This description is then encapsulated with the set of evidence items to form a traceable evidence package.

[0047] As a preferred embodiment of the multimodal fusion-based early warning method for cross-contamination of raw and cooked food in the kitchen according to the present invention, the specific steps for triggering real-time early warning of cross-contamination in the kitchen scenario are as follows:

[0048] Extract pollution path descriptions from traceable evidence packages to generate early warning event records;

[0049] The warning event record is combined with the corresponding key frame evidence to form a warning release message, which is then sent to the kitchen audio-visual terminal and display terminal to generate a real-time warning event record of cross-contamination.

[0050] The records of real-time early warning events of cross-contamination are archived along with traceable evidence packages, and a search index containing time, work area, and work entity identification is generated to form a risk profile.

[0051] Secondly, this invention provides a multimodal fusion-based early warning system for cross-contamination between raw and cooked food ingredients in a kitchen, including:

[0052] The acquisition and registration module acquires visible light video, thermal infrared video, and depth / binocular data in the kitchen operation area and completes spatiotemporal registration to determine target candidates for kitchen operation entities and generate a multimodal object candidate set.

[0053] The relationship recognition module performs relationship discrimination on the multimodal object candidate set according to the open vocabulary video relationship pattern recognition rules, and outputs triple event streams and corresponding object identity trajectory tables;

[0054] The process assembly module determines the stage sequence of the kitchen operation process based on the event flow of relation triples and the object identity trajectory table, and completes the assembly of contamination events to form a cross-contamination evidence chain.

[0055] The early warning and archiving module performs consistency verification on the cross-contamination evidence chain, forms a traceable evidence package, triggers real-time early warning of cross-contamination in the kitchen scenario, and generates a risk profile.

[0056] The beneficial effects of this invention are as follows: By utilizing a unified time index and table plane coordinate system of visible light video, thermal infrared video, and depth / binocular data, the invention achieves a homologous expression of object short trajectories and 3D pose estimation results, enabling the spatial migration of processing tools, carrying utensils, personnel hands, and food objects between raw food areas, cooked food areas, washing and disinfection areas, and serving outlets to have a computable basis; by employing open-vocabulary video relational pattern recognition to output triplet event streams and object identity trajectory tables, the invention ensures that interactive relationships such as "contact, carrying, passing, picking up, and placing" are stably implemented in the form of event fragments, thereby supporting subsequent stage sequence division and contamination candidate link retrieval; by generating consistency markers through cross-modal consistency verification and encapsulating the evidence item set and link description into a traceable evidence package, the invention enables early warning issuance to simultaneously provide interpretable contamination paths and searchable risk files, significantly improving the reliability and traceability of kitchen cross-contamination early warning. Attached Figure Description

[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 is a flowchart of a method for early warning of cross-contamination between raw and cooked food ingredients in the kitchen based on multimodal fusion.

[0059] Figure 2 is a flowchart of multimodal data acquisition and processing.

[0060] Figure 3 is a flowchart of multimodal object recognition and tracking.

[0061] Figure 4 is a flowchart of cross-contamination early warning and evidence chain generation. Detailed Implementation

[0062] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0063] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0064] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0065] Referring to Figures 1-4, an embodiment of the present invention is provided, which offers a method for early warning of cross-contamination between raw and cooked food ingredients in a kitchen based on multimodal fusion, comprising the following steps:

[0066] S1. Acquire visible light video, thermal infrared video and depth / binocular data in the kitchen operation area and complete spatiotemporal registration to determine target candidates for kitchen operation entities and generate a multimodal object candidate set.

[0067] S1.1. Visible light cameras, thermal infrared cameras, and depth / binocular cameras are respectively deployed in the raw food area, cooked food area, washing and disinfection area, and food outlet of the kitchen operation area. Visible light video, thermal infrared video, and depth / binocular data are collected during the same operation period, and a timestamp is recorded for each frame to obtain a visible light video sequence, thermal infrared video sequence, and depth / binocular data sequence containing timestamps.

[0068] The visible light video, thermal infrared video, and depth / stereo data are time-stamp aligned. The correspondence between visible light frames, thermal infrared frames, and depth / stereo frames is established for each moment according to the nearest neighbor matching method of timestamp, and a unified time index is generated.

[0069] S1.2. After timestamp alignment, multi-view calibration images are acquired using a calibration board under the views of a visible light camera, a thermal infrared camera, and a depth / binocular camera. Geometric feature points such as corner points or dots of the calibration board are extracted from each calibration image, and the correspondence between the feature points and the plane coordinates and pixel coordinates of the calibration board is established. Based on the correspondence, the camera intrinsic parameter calibration results of the visible light camera, thermal infrared camera, and depth / binocular camera are calculated. The camera intrinsic parameter calibration results include focal length parameters, principal point parameters, and distortion parameters.

[0070] Distortion correction is performed on visible light frames with a unified time index using the camera intrinsic calibration results of a visible light camera, on thermal infrared frames with a unified time index using the camera intrinsic calibration results of a thermal infrared camera, and on depth / binocular frames with a unified time index using the camera intrinsic calibration results of a depth / binocular camera. The distortion correction is performed by inversely mapping the pixel coordinates to the ideal imaging plane and then interpolating and resampling to obtain a corrected image sequence.

[0071] Calibration images of the same calibration board are acquired from the perspectives of a visible light camera, a thermal infrared camera, and a depth / binocular camera. Feature points of the calibration board are extracted from each image, and a one-to-one correspondence between these feature points and the pixel coordinate system is established. Utilizing the geometric constraint that the calibration board is a planar target, the pose from the calibration board coordinate system to the camera coordinate system is calculated for each perspective. This ensures that the projected positions of the calibration board feature points after camera imaging are consistent with the observed pixel coordinates. The imaging constraint relationship indicates that the projected coordinates of the calibration board feature points and the pixel coordinates of the feature points detected in the calibration images satisfy the same imaging model. This imaging model is described by both the camera intrinsic calibration results and the camera extrinsic calibration results. The camera extrinsic calibration results include rotation matrices and translation vectors, which are used to establish the coordinate mapping relationship between visible light, thermal infrared, and depth / binocular data.

[0072] The camera extrinsic parameter calibration results satisfy the rigid body transformation relationship between the calibration plate coordinate system and the camera coordinate system, specifically:

[0073] ;

[0074] in, This represents the three-dimensional coordinates of the feature points on the calibration plate in the calibration plate coordinate system. This represents the three-dimensional coordinates of the feature points on the calibration board in the camera coordinate system. This represents the rotation matrix in the camera extrinsic calibration results, used to describe the coordinate axis orientation relationships. This represents the translation vector in the camera extrinsic calibration results, used to describe the offset of the coordinate origin position.

[0075] S1.3. Extract the platform point cloud from the depth / stereo data and perform plane fitting to obtain the operation plane extraction result. The operation plane extraction result includes the platform plane normal vector and the spatial position of a point on the platform plane. Determine the platform plane coordinate system based on the operation plane extraction result and unify the visible light frame, thermal infrared frame and depth / stereo frame to the platform plane coordinate system to generate multimodal aligned data with unified coordinates and unified time index.

[0076] To further explain, plane fitting refers to first selecting points in the depth / stereo data that have valid depth and fall within a preset height range (obtained by expanding vertically after determining the center height of the platform by the peak position of the height histogram of the platform point cloud in the depth / stereo data, for example, taking a depth interval of 20 centimeters above and below the nominal height of the platform) as a candidate point set for the platform point cloud. Outlier removal is then performed on the candidate point set to eliminate flying points and reflection noise points. Finally, plane fitting based on random sampling consistency is performed. Random sampling consistency plane fitting includes extracting three non-common points from the candidate point set of the platform point cloud. The process involves calculating the plane parameters using line points, calculating the distance from each point in the candidate point set of the platform point cloud to the plane, and counting the number of interior points that meet the distance threshold (determined by the ranging accuracy index of depth / binocular data and the statistical results of the local plane residual of the platform point cloud, set to 5mm). After iteration, the plane parameter with the largest number of interior points is selected. Then, the least squares method is used to refit the plane parameters on the set of interior points, and the platform plane normal vector and the spatial position of a point on the platform plane are updated to obtain the operation plane extraction result, which includes the platform plane normal vector and the spatial position of a point on the platform plane.

[0077] The operation plane extraction result satisfies the table plane constraint relationship, specifically:

[0078] ;

[0079] in, This represents the three-dimensional coordinates of any point within the point cloud of the platform. Represents the three-dimensional coordinates of a point on the platform plane. This represents the normal vector of the platform plane. The normal vector of the platform plane and a point on the platform plane are used to determine the coordinate system of the platform plane and are used for subsequent multimodal coordinate unification and expression of 3D pose estimation results.

[0080] On visible light frames, thermal infrared frames, and depth / binocular frames carrying multimodal aligned data with unified coordinates and unified time indexes, target detection and instance segmentation, target segmentation and contour extraction, and depth clustering and 3D boundary extraction are performed respectively to obtain visible light candidate sets, thermal infrared candidate sets, and depth candidate sets.

[0081] To further explain, object detection and instance segmentation refer to the process of normalizing the size and pixels of the visible light frame to obtain candidate boxes and class scores, as well as pixel-level segmentation masks corresponding to the candidate boxes. Non-maximum suppression is performed on the candidate boxes to remove overlapping and redundant candidates. The visible light candidate set is output by combining the class scores and the segmentation masks.

[0082] Target segmentation and contour extraction refers to normalizing the temperature or grayscale values ​​of the thermal infrared frame, using adaptive threshold segmentation to obtain the foreground region, performing morphological opening and closing operations to remove small noise and fill holes, using connected component analysis to obtain candidate regions and extract the outer contours of the candidate regions, and outputting a thermal infrared candidate set.

[0083] Deep clustering and 3D boundary calculation involves converting depth / binocular frames into point clouds and projecting them onto a platform plane coordinate system. Euclidean distance clustering is then used to segment the point clouds into several 3D point clusters. For each 3D point cluster, a 3D bounding box is calculated to obtain the 3D boundary extraction result. The 3D bounding box calculation process includes performing principal component analysis on the principal directions of the 3D point clusters in the platform plane coordinate system to obtain the orientation direction, and then calculating the boundary range of the point clusters in the orientation direction to output a depth candidate set.

[0084] Cross-modal association is performed on the visible light candidate set, the thermal infrared candidate set, and the depth candidate set. The cross-modal association is performed by projecting the thermal infrared contour and the 3D depth boundary onto the visible light coordinates using a cross-modal mapping matrix and performing spatial overlap determination on the region segmented by the visible light instance. Candidate combinations that meet the spatial overlap determination are formed into entity-level matching pairs.

[0085] S1.4. Perform short-term tracking of entity-level matching pairs along the time dimension. Short-term tracking includes establishing associations based on the spatial continuity and appearance similarity of entity-level matching pairs in adjacent frames and generating short object trajectories, while assigning trajectory identifiers to each object short trajectory.

[0086] The 3D center point and orientation information of the object are calculated using the 3D boundary of the depth candidate set and the table plane coordinate system to obtain the 3D pose estimation result. The 3D pose estimation result is bound to the trajectory identifier and the object short trajectory. The visible light candidate set, thermal infrared candidate set, depth candidate set, entity-level matching pair, object short trajectory, trajectory identifier and 3D pose estimation result are summarized to output a multimodal object candidate set containing target candidates such as processing tools, bearing equipment, operating plane area, personnel hands and food objects.

[0087] S2. Perform relation discrimination on the multimodal object candidate set according to the open vocabulary video relation pattern recognition rules, and output the triple event stream and the corresponding object identity trajectory table.

[0088] S2.1. Construct an entity category list based on kitchen operation entities and generate kitchen operation entity phrases. The entity category list should at least cover processing tools, carrying utensils, operating plane areas, personnel hands, and food objects. Construct a kitchen interaction relation lexicon based on kitchen operation process semantics and cover contact, carrying, passing, picking up, and placing relation categories. Combine the kitchen operation entity phrases with the kitchen interaction relation lexicon to generate a relation text sequence. Perform text encoding on the relation text sequence to generate a relation semantic prototype library. Combine the relation semantic prototype library with the relation text sequence to form a relation prompt set.

[0089] The short object trajectories in the multimodal object candidate set are summarized according to the trajectory identifier. The trajectory temporal features are extracted for each short object trajectory and the entity category label is determined. The trajectory temporal features include the center point sequence, orientation sequence and visible light instance segmentation mask sequence of the short object trajectory. The entity category label is jointly determined by the category score of the visible light candidate set and the 3D boundary shape features of the depth candidate set. The entity category label is used to divide the short object trajectory into the subject trajectory set and the object trajectory set and maintain the consistency of the trajectory identifier in the subject trajectory set and the object trajectory set.

[0090] S2.2. Calculate the time overlap segment of any two short trajectories of objects in the subject trajectory set and object trajectory set and generate an overlap time window. Within each overlap time window, calculate the spatial distance between the center point of the subject trajectory and the center point of the object trajectory in the table plane coordinate system and filter the spatially adjacent segments. Within the overlap time window, calculate the spatial distance between the center point of the subject trajectory and the center point of the object trajectory frame by frame, and merge the frame intervals whose spatial distance is continuously less than the preset adjacent distance threshold (determined by the statistical results of the typical contact distance between the processing tool and the food object in the table plane coordinate system and taken within the three-dimensional boundary size range of the kitchen operation entity, such as 5cm) into the same spatially adjacent segment. Combine the overlap time window and the spatially adjacent segments to generate a set of relational spatiotemporal segments.

[0091] To further explain, the expression for spatial distance is:

[0092] ;

[0093] in, This represents the spatial distance between the center point of the subject's trajectory and the center point of the object's trajectory. This represents the coordinate vector of the center point of the main trajectory in the coordinate system of the platform plane. This represents the coordinate vector of the center point of the object's trajectory in the coordinate system of the table plane. This represents Euclidean norm operations.

[0094] Visible light appearance features, thermal infrared contour features, and 3D pose change features are extracted segment by segment from the spatiotemporal fragment set of the relation and fused to form a relation feature sequence. The visible light appearance features consist of the texture and shape description of the subject instance segmentation mask region and the object instance segmentation mask region of the corresponding frame of the spatiotemporal fragment set of the relation. The thermal infrared contour features consist of the geometric description of the outer contour of the thermal infrared candidate set. The 3D pose change features consist of the displacement and orientation change of the center point of the 3D pose estimation result within the corresponding time window of the spatiotemporal fragment set of the relation.

[0095] S2.3. Perform cross-modal contrast matching on the relation feature sequence and relation cue set in the visual-language joint model. Cross-modal contrast matching includes visual feature encoding of the relation feature sequence, text feature encoding of the relation text sequence, and calculating the similarity between visual features and text features in the same feature space, and outputting the relation score sequence.

[0096] To further explain, the expression for calculating similarity is:

[0097] ;

[0098] in, Indicates the similarity between visual features and text features. This represents the feature vector obtained by visual feature encoding of the relational feature sequence. The feature vector represents the relational text sequence obtained by text feature encoding.

[0099] Perform temporal consistency decoding on the relation score sequence and determine the start and end range of the relation. Temporal consistency decoding includes performing temporal smoothing on the relation score sequences of adjacent frames and performing continuous segment merging. The time index of the first frame of the continuous segment is determined as the relation start time index, and the time index of the last frame of the continuous segment is determined as the relation end time index, forming relation triplet event segments.

[0100] The event fragments of the relation triplet are associated with the corresponding subject short trajectory identifiers and object short trajectory identifiers. The associated registration content includes the subject short trajectory identifier, relation category, object short trajectory identifier, relation start time index and relation end time index. The associated registration content is summarized in the order of time index to generate the triplet event flow, and the subject short trajectory identifiers and object short trajectory identifiers are summarized in the order of trajectory identifiers to generate the object identity trajectory table.

[0101] S3. Based on the event flow of relation triples and the object identity trajectory table, determine the stage sequence of the back kitchen operation process and complete the assembly of contamination events to form a cross-contamination evidence chain.

[0102] S3.1. Read the main short trajectory identifiers corresponding to the processing tools, the main short trajectory identifiers corresponding to the personnel's hands, and the main short trajectory identifiers corresponding to the food objects from the object identity trajectory table. In the multimodal object candidate set, retrieve the object short trajectory and 3D pose estimation results that are consistent with the main short trajectory identifiers. Concatenate the center point sequence and orientation sequence of the object short trajectory according to a unified time index to form the tool trajectory, hand trajectory, and food trajectory. At the same time, map the tool trajectory, hand trajectory, and food trajectory to the work area range corresponding to the raw food area, cooked food area, washing and disinfection area, and food outlet according to the table plane coordinate system. Output a set of motion fragments with work area annotations.

[0103] S3.2. Align the action fragment set with the relation triplet event flow according to a unified time index. When the relation start time index and relation end time index in the relation triplet event flow fall within the time range of the action fragment set, establish a time association between the relation triplet event fragment and the action fragment set. The time association result is divided into stages according to the operation semantics to obtain a stage sequence. The operation semantics are jointly determined by the relation category and the work area label. The cutting and preparation stage corresponds to the processing tool and the food object having a contact relationship and the work area is labeled as the raw food area or the cooked food area. The plating stage corresponds to the carrying vessel and the food object having a carrying relationship and the work area is labeled as the cooked food area or the serving port. The washing and disinfection stage corresponds to the processing tool and the operation plane area or the carrying vessel having a contact relationship and the work area is labeled as the washing and disinfection area. The temporary storage stage corresponds to the carrying vessel and the operation plane area having a placement relationship and the work area is labeled as the raw food area or the cooked food area.

[0104] S3.3. Time-align the stage sequence with the relation triplet event stream. Each relation triplet event fragment in the relation triplet event stream is assigned to the corresponding stage name according to the relation start time index and relation end time index. The subject short trajectory identifier, object short trajectory identifier and entity category label in the object identity trajectory table are used to establish entity association. The entity association result constructs a stage relationship graph with stage name as node and relation triplet event fragment as edge.

[0105] S3.4. Event fragments of relation triplets in the stage relationship diagram that belong to the raw food area and contain food objects are labeled as raw food-related events. Event fragments of relation triplets in the stage relationship diagram that belong to the cooked food area or serving area and contain food objects are labeled as cooked food-related events. Raw food-related events and cooked food-related events are searched for entity consistency based on the subject short trajectory identifier and the object short trajectory identifier. The entity consistency search detects whether the processing tools, carrying equipment, and personnel's hands appear with the same subject short trajectory identifier or the same object short trajectory identifier in the raw food-related events and cooked food-related events. The cross-area migration order of the same subject short trajectory identifier or the same object short trajectory identifier before the disinfection stage is searched in the stage sequence. The contamination candidate links are output and the link start time index, link end time index, subject short trajectory identifier, object short trajectory identifier, and work area label of the contamination candidate links are recorded.

[0106] S3.5. Sort the contamination candidate links according to the link start time index and match them one by one with the edge set of the stage relationship graph. Based on the matching results, extract the relationship category sequence, stage name sequence and work area label sequence corresponding to the contamination candidate links. Retrieve the short trajectory and three-dimensional pose estimation results of the objects involved in the contamination candidate links in the multimodal object candidate set to complete the spatial location. Assemble the relationship category sequence, stage name sequence, time range and spatial location of the contamination candidate links in chronological order to form a cross-contamination evidence chain.

[0107] S4. Perform consistency verification on the cross-contamination evidence chain to form a traceable evidence package, trigger real-time early warning of cross-contamination in the kitchen scenario, and generate a risk profile.

[0108] S4.1. Extract the link start time index, link end time index, relationship category sequence and work area label sequence of the cross-contamination evidence chain candidate links. Extract key nodes according to the boundary positions of the link start time index, link end time index and relationship category sequence, and summarize the time indexes corresponding to the key nodes to generate key frame index.

[0109] It should be noted that key nodes refer to the relationship triplet event fragments corresponding to the start position, end position, and position where the relationship category changes in the cross-contamination evidence chain of candidate links.

[0110] The visible light candidate set corresponding to the visible light frame indicated by the keyframe index is extracted as visible light evidence, the instance segmentation mask and candidate box region corresponding to the visible light candidate set are extracted as visible light evidence, the thermal infrared candidate set corresponding to the thermal infrared frame indicated by the keyframe index is extracted as thermal infrared evidence, the depth candidate set corresponding to the 3D bounding box and point cloud fragments are extracted as depth / stereo evidence from the depth / stereo frame indicated by the keyframe index, and the thermal infrared evidence and depth / stereo evidence are projected onto the visible light frame coordinate system using a cross-modal mapping matrix to form a projected contour region.

[0111] S4.2. When performing cross-modal consistency verification on the projected contour region and visible light evidence, for the subject short trajectory identifier and object short trajectory identifier indicated by the object identity trajectory table, calculate the overlap between the instance segmentation mask region and the projected contour region of the visible light evidence in the corresponding frame of the key frame index, and associate the overlap result with the subject short trajectory identifier, object short trajectory identifier, relationship category and key frame index to generate a consistency mark.

[0112] The expression for overlap is:

[0113] ;

[0114] in, Indicates the degree of overlap. This represents the set of pixels in the instance segmentation mask region of visible light evidence. This represents the set of pixels representing the projected contour region obtained from thermal infrared evidence or depth / binocular evidence projection. The intersection operation represents the set of pixels. This represents the union operation of pixel sets.

[0115] When performing the association and aggregation of consistency markers with object identity trajectory tables, triple event flows, and stage sequences, keyframe indices are matched according to the subject short trajectory identifier, object short trajectory identifier, relationship start time index, and relationship end time index of the triple event flow. Similarly, triple event flows are matched according to the stage start time index and stage end time index of the stage sequence. The matching results are then merged with the consistency markers to generate a set of evidence entries. This set of evidence entries includes subject short trajectory identifiers, object short trajectory identifiers, entity category labels, relationship categories, stage names, work area labels, keyframe indices, and corresponding references to visible light evidence, thermal infrared evidence, and depth / binocular evidence.

[0116] S4.3. When generating a link description by combining the time sequence of key nodes in the cross-contamination evidence chain with the connection relationship of adjacent nodes, organize the pollution propagation path text according to the relationship category sequence, stage name sequence, and work area labeling sequence of the cross-contamination evidence chain. The pollution propagation path text and the set of evidence items are jointly packaged to form a traceable evidence package.

[0117] The link description is extracted from the traceable evidence package to form a pollution path description. The pollution path description is combined with the visible light evidence, thermal infrared evidence and depth / binocular evidence corresponding to the key frame index to generate an early warning event record. The early warning event record is encapsulated into an early warning release message and sent to the kitchen audio-visual terminal and display terminal to form a real-time early warning event record of cross-contamination.

[0118] When archiving the real-time early warning event records and traceable evidence packages for cross-contamination, a retrieval index is generated according to the time index, work area label, and work entity identifier in the object identity trajectory table in the real-time early warning event records for cross-contamination. The retrieval index, together with the real-time early warning event records for cross-contamination and the traceable evidence packages, constitutes a risk file.

[0119] This embodiment also provides a multimodal fusion-based early warning system for cross-contamination between raw and cooked food in the kitchen, including:

[0120] The acquisition and registration module acquires visible light video, thermal infrared video, and depth / binocular data in the kitchen operation area and completes spatiotemporal registration to determine target candidates for kitchen operation entities and generate a multimodal object candidate set.

[0121] The relationship recognition module performs relationship discrimination on the multimodal object candidate set according to the open vocabulary video relationship pattern recognition rules, and outputs triple event streams and corresponding object identity trajectory tables;

[0122] The process assembly module determines the stage sequence of the kitchen operation process based on the event flow of relation triples and the object identity trajectory table, and completes the assembly of contamination events to form a cross-contamination evidence chain.

[0123] The early warning and archiving module performs consistency verification on the cross-contamination evidence chain, forms a traceable evidence package, triggers real-time early warning of cross-contamination in the kitchen scenario, and generates a risk profile.

[0124] In summary, this invention utilizes a unified temporal index and table plane coordinate system for visible light video, thermal infrared video, and depth / binocular data to achieve a homologous representation of object short trajectories and 3D pose estimation results. This provides a computational basis for the spatial migration of processing tools, carrying containers, personnel hands, and food objects between raw food areas, cooked food areas, washing and disinfection areas, and the food outlet. It employs open-vocabulary video relational pattern recognition to output triplet event streams and object identity trajectory tables, ensuring that interactive relationships such as "contact, carrying, passing, picking up, and placing" are stably represented as event fragments, thus supporting subsequent stage sequence segmentation and contamination candidate link retrieval. Furthermore, it generates consistency markers through cross-modal consistency verification and encapsulates the evidence item set and link description into a traceable evidence package, enabling early warning issuance to simultaneously provide interpretable contamination paths and searchable risk profiles, significantly improving the reliability and traceability of kitchen cross-contamination early warnings.

[0125] Example 2, referring to Table 1, is the second embodiment of the present invention. To further verify the technical solution of the present invention, experimental simulation data of a method for early warning of cross-contamination between raw and cooked food ingredients in the kitchen based on multimodal fusion are given.

[0126] In a centralized food preparation kitchen environment, the raw food area, cooked food area, washing and disinfection area, and food outlet were selected as the kitchen operation areas. Visible light cameras, thermal infrared cameras, and depth / binocular cameras were deployed to collect data during the same operation period. The visible light cameras collected visible light video with a resolution of 1920×1080 and a frame rate of 30 frames per second. The thermal infrared cameras collected thermal infrared video with a resolution of 640×512 and a frame rate of 30 frames per second. The depth / binocular cameras collected depth / binocular data with a resolution of 1280×720 and recorded a timestamp for each frame. The correspondence between visible light frames, thermal infrared frames, and depth / binocular frames was established using a timestamp nearest neighbor matching method. A unified temporal index is generated based on the relationship. Subsequently, a calibration board is used to acquire multi-view calibration images and extract corner or round feature points. The camera intrinsic calibration results are solved by the correspondence between planar coordinates and pixel coordinates, and distortion correction is performed. The camera extrinsic calibration results are solved by the one-to-one correspondence of feature points and planar geometric constraints, and a cross-modal mapping matrix is ​​generated. Then, the platform point cloud is extracted from the depth / binocular data, and the operation plane extraction results are obtained by random sampling consistency plane fitting and least squares refitting. The platform plane coordinate system is determined by the operation plane extraction results, and the multimodal coordinate unification is completed, generating multimodal aligned data with unified coordinates and a unified temporal index.

[0127] Based on multimodal alignment data, target detection and instance segmentation are performed in the visible light frame to obtain a visible light candidate set, target segmentation and contour extraction are performed in the thermal infrared frame to obtain a thermal infrared candidate set, and depth clustering and 3D boundary extraction are performed in the depth / binocular frame to obtain a depth candidate set. Cross-modal association is completed using a cross-modal mapping matrix to form entity-level matching pairs. Subsequently, short-time tracking is performed along the time dimension to generate short object trajectories and assign trajectory labels. The 3D pose estimation results are calculated using the 3D boundary of the depth candidate set and the platform plane coordinate system. Finally, the results are summarized to obtain a multimodal object candidate set.

[0128] In the relation identification stage, kitchen operation entity phrases are generated based on kitchen operation entities, and a kitchen interaction relation vocabulary is generated based on kitchen operation process semantics. These are combined to obtain a relation text sequence, and text encoding is performed to generate a relation semantic prototype library. This library is then combined with the relation text sequence to form a relation hint set. Trajectory temporal features are extracted from short object trajectories, and entity category labeling is completed. The overlap time window and spatial proximity segments of the subject trajectory and object trajectory are calculated to form a relation spatiotemporal segment set. Visible light appearance features, thermal infrared contour features, and three-dimensional pose change features are extracted and fused to obtain a relation feature sequence. Cross-modal contrast matching is performed in the visual-language joint model to output a relation score sequence, which is then further processed for temporal consistency. The code forms a relational triplet event fragment and registers it to generate a triplet event flow and an object identity trajectory table. Finally, based on the object identity trajectory table and the triplet event flow, the stage sequence is determined and a stage relationship graph is constructed. Entity consistency retrieval is performed on raw food-related events and cooked food-related events to obtain contamination candidate links and assemble cross-contamination evidence chains. Keyframe indexes are extracted for cross-contamination evidence chains and cross-modal consistency verification is completed to generate consistency markers. The evidence entries are associated and summarized to form a set of evidence items and packaged into a traceable evidence package. Contamination path descriptions are extracted to form early warning event records and generate cross-contamination real-time early warning event records. At the same time, a retrieval index containing time, work area, and work entity identifier is archived to form a risk file.

[0129] The details are shown in Table 1 below:

[0130] Table 1. Experimental Data on Multimodal Validation of Cross-Contamination in the Kitchen

[0131] Parameters / Units This Invention - Multimodal Open Lexical Relationship + Consistency Check Comparison 1 - Visible Light Instance Segmentation Comparison 2 - Multimodal No Spatiotemporal Registration Comparison 3 - Multimodal Registration + Closed Relationship Classification Comparison 4 - Open Lexical Relationship No Consistency Check Comparison 5 - Rule Link Stitching Method Time Alignment Error (ms) 6.4 7.1 18.6 6.8 6.67 Cross-modal Reprojection Error (px) 1.3 4.9 5.6 1.6 1.5 2.4 Platform Plane Fitting Residual (mm) 1.8 6.7 21.9 1.8 2.1 Multimodal Object Candidate mAP (%) ) 91.678.483.790.191.288.5 Relationship triplet identification F1 (%) 86.255.159.374.885.162.4 Cross-contamination evidence chain recall rate (%) 92.460.265.578.990.671.8 Cross-contamination early warning false alarm rate (times / hour) 0.62.82.11.41.71.9 Early warning event generation delay (ms) 320410460350300520 Traceable evidence package integrity (0-100) 954251797258 surface

[0132] As can be seen from the "cross-modal reprojection error (px)" and "table plane fitting residual (mm)," this invention forms a cross-modal mapping matrix and completes multimodal coordinate unification by using camera intrinsic parameter calibration results, camera extrinsic parameter calibration results, and operation plane extraction results. This significantly improves the consistency of cross-modal space. The cross-modal reprojection error decreased from 4.9px in Comparison 1 to 1.3px, and the table plane fitting residual decreased from 6.7mm in Comparison 1 to 1.8mm. This indicates that the spatial reference of the multimodal alignment data is more stable, fundamentally improving the consistent representation ability of subsequent object short trajectory and 3D pose estimation results.

[0133] As can be seen from the "F1 score of relation triple identification (%)" and "recall rate of cross-contamination evidence chain (%)", this invention uses a visual-language joint model to perform cross-modal comparison matching on relation feature sequences and relation cue sets, which improves relation identification from the traditional "only identifying object categories" to "directly identifying object interaction relationships". The F1 score of relation triple identification reaches 86.2%, which is significantly higher than 74.8% of comparison 3 and 62.4% of comparison 5. At the same time, the recall rate of cross-contamination evidence chain reaches 92.4%, which is 32.2 percentage points higher than comparison 1. This shows that open vocabulary video relation pattern recognition has a stronger relation expression ability in complex interactive scenarios in the kitchen and can stably support the generation and assembly of contamination candidate links.

[0134] As can be seen from the "false alarm rate of cross-contamination early warning (times / hour)" and the "completeness of traceable evidence package (0-100)", this invention, through keyframe indexing, cross-modal consistency verification and consistency marking, unifies visible light evidence, thermal infrared evidence and depth / binocular evidence into a set of evidence items and forms a traceable evidence package. The false alarm rate is reduced to 0.6 times / hour, which is significantly better than the 1.7 times / hour of comparison 4. At the same time, the completeness of the evidence package reaches 95, which is higher than the 72 of comparison 4. This shows that the consistency verification not only enhances the credibility of the early warning, but also forms a structured and searchable risk file, which can cover the needs of regulatory traceability and on-site review, and realize an interpretable closed-loop output of "identifying objects - identifying relationships - identifying pollution chains".

[0135] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for early warning of cross-contamination between raw and cooked food in a kitchen based on multimodal fusion, characterized in that: This includes acquiring visible light video, thermal infrared video, and depth / binocular data in the kitchen work area and completing spatiotemporal registration, determining target candidates for kitchen work entities, and generating a multimodal object candidate set; Relationship discrimination is performed on the multimodal object candidate set according to the open vocabulary video relationship pattern recognition rules, and the triple event stream and the corresponding object identity trajectory table are output. Based on the event flow of relation triples and the object identity trajectory table, the stage sequence of the kitchen operation process is determined and the contamination event is assembled to form a cross-contamination evidence chain. Consistency verification is performed on the cross-contamination evidence chain to form a traceable evidence package, triggering a real-time early warning of cross-contamination in the kitchen scenario and generating a risk profile.

2. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 1, characterized in that: The kitchen work area includes a raw food area, a cooked food area, a washing and disinfection area, and a food serving outlet; the kitchen work entities include processing tools, carrying containers, operating surfaces, personnel hands, and food items.

3. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 1, characterized in that: The specific steps for completing the spatiotemporal registration are as follows: After aligning the visible light video, thermal infrared video, and depth / binocular data with timestamps, a calibration board is used to acquire multi-view calibration images and calculate the camera intrinsic parameter calibration results of the visible light camera, thermal infrared camera, and depth / binocular camera. The camera intrinsic parameter calibration results are used to complete the distortion correction of each viewpoint and generate a corrected image sequence. The corresponding feature points of the same calibration board under each camera viewpoint are acquired and the camera extrinsic parameter calibration results are calculated. The camera extrinsic parameter calibration results are used to establish the coordinate mapping relationship between visible light, thermal infrared, and depth / binocular data and generate a cross-modal mapping matrix. Extract platform point clouds from depth / stereo data and perform plane fitting to obtain operational plane extraction results. Use the operational plane extraction results to determine the platform plane coordinate system and complete the unification of multimodal coordinates. Output multimodal aligned data with unified coordinates and unified time index.

4. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 1, characterized in that: The specific steps for generating the multimodal object candidate set are as follows: On the visible light frame, thermal infrared frame, and depth / binocular frame of the multimodal aligned data, target detection and instance segmentation, target segmentation and contour extraction, and depth clustering and 3D boundary extraction are performed respectively to obtain the visible light candidate set, thermal infrared candidate set, and depth candidate set; the visible light candidate set, thermal infrared candidate set, and depth candidate set are cross-modal associated, and short-term tracking is performed along the time dimension to generate short object trajectories and assign trajectory identifiers; the 3D center point and orientation information of the object are calculated using the 3D boundary of the depth candidate set and the table plane coordinate system, and bound to the trajectory identifiers of the object's short trajectory to output the multimodal object candidate set.

5. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 1, characterized in that: The specific steps for determining relationships in a multimodal object candidate set according to the open vocabulary video relationship pattern recognition rules are as follows: First, obtain the kitchen interaction relationship vocabulary and kitchen operation entity phrases to form a relationship hint set. Second, extract the trajectory temporal features of each object short trajectory in the multimodal object candidate set and determine the entity category label. Simultaneously, divide the object short trajectories into a subject trajectory set and an object trajectory set according to the entity category label. Third, calculate the temporal overlap and spatial proximity segments of any two object short trajectories in the subject trajectory set and object trajectory set to generate a relationship spatiotemporal segment set. Fourth, extract visible light appearance features, thermal infrared contour features, and three-dimensional pose change features from the relationship spatiotemporal segment set and fuse them to form a relationship feature sequence. Fifth, perform cross-modal comparison matching between the relationship feature sequence and the relationship hint set in a visual-language joint model to generate a relationship score sequence. Sixth, perform temporal consistency decoding on the relationship score sequence and determine the start and end range of the relationship to form a relationship triplet event segment. Seventh, associate and register the relationship triplet event segment with the corresponding subject short trajectory identifier and object short trajectory identifier to generate a triplet event stream and a corresponding object identity trajectory table.

6. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 5, characterized in that: The specific steps for obtaining the kitchen interaction relation lexicon and the kitchen operation entity phrase library are as follows: construct an entity category list based on kitchen operation entities, generate kitchen operation entity phrases; construct a kitchen interaction relation lexicon based on the semantics of kitchen operation processes. The entity phrases of kitchen operations are combined with the kitchen interaction relation vocabulary to generate a relation text sequence; the relation text sequence is text encoded to generate a relation semantic prototype library; the relation semantic prototype library is combined with the relation text sequence to form a relation hint set.

7. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 1, characterized in that: The specific steps for determining the stage sequence of the kitchen operation process and assembling the contamination event are as follows: extract tool trajectory, hand trajectory and food trajectory from the object identity trajectory table to obtain a set of action segments; divide the set of action segments into stages according to the operation semantics to obtain the stage sequence; Align the phase sequence with the triplet event flow in time to form a phase relationship graph; Cross-stage link search is performed on raw food-related events and cooked food-related events in the stage relationship diagram to obtain contamination candidate links; the contamination candidate links are assembled according to time sequence and spatial location to form a cross-contamination evidence chain.

8. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 1, characterized in that: The specific steps for performing consistency verification on the cross-contamination evidence chain are as follows: extract key nodes from the cross-contamination evidence chain and generate key frame indexes; perform cross-modal consistency verification on the visible light evidence, thermal infrared evidence and depth / binocular evidence indicated by the key frame indexes and generate consistency markers. Consistency markers, object identity trajectory tables, triplet event flows, and stage sequences are correlated and summarized to generate a set of evidence items. Link descriptions are generated according to the time sequence of key nodes in the cross-contamination evidence chain and the connection relationship between adjacent nodes, and then encapsulated with the set of evidence items to form a traceable evidence package.

9. The method for early warning of cross-contamination between raw and cooked food in the kitchen based on multimodal fusion as described in claim 1, characterized in that: The specific steps for triggering the real-time cross-contamination early warning in the kitchen scenario are as follows: extract the contamination path description from the traceable evidence package to generate an early warning event record; combine the early warning event record with the corresponding keyframe evidence to form an early warning release message, and send it to the kitchen audio-visual terminal and display terminal to generate a real-time cross-contamination early warning event record; The records of real-time early warning events of cross-contamination are archived along with traceable evidence packages, and a search index containing time, work area, and work entity identification is generated to form a risk profile.

10. A multimodal fusion-based early warning system for cross-contamination of raw and cooked food ingredients in a kitchen, based on the multimodal fusion-based early warning method for cross-contamination of raw and cooked food ingredients in a kitchen as described in any one of claims 1 to 9, characterized in that: This includes a data acquisition and registration module, which acquires visible light video, thermal infrared video, and depth / binocular data in the kitchen work area and completes spatiotemporal registration, determines target candidates for kitchen work entities, and generates a multimodal object candidate set; The relationship recognition module performs relationship discrimination on the multimodal object candidate set according to the open vocabulary video relationship pattern recognition rules, and outputs triple event streams and corresponding object identity trajectory tables; The process assembly module determines the stage sequence of the kitchen operation process based on the event flow of relation triples and the object identity trajectory table, and completes the assembly of contamination events to form a cross-contamination evidence chain. The early warning and archiving module performs consistency verification on the cross-contamination evidence chain, forms a traceable evidence package, triggers real-time early warning of cross-contamination in the kitchen scenario, and generates a risk profile.