Dietary intake monitoring method and apparatus, computer device, and storage medium

By using PTZ camera equipment and a pre-set segmentation network to automatically identify the food categories and remaining quantities of kindergarten children, the problem of low efficiency in monitoring food intake in kindergartens has been solved, achieving unobtrusive monitoring of food intake and efficient food quantity analysis.

CN115497035BActive Publication Date: 2026-05-15GUANGZHOU PAIKEPUSHI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU PAIKEPUSHI INFORMATION TECH CO LTD
Filing Date
2022-07-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, the monitoring of children's dietary intake in kindergartens is inefficient, and traditional methods cannot obtain specific dietary information and increase labor costs.

Method used

By acquiring images of the dining scene using a PTZ camera device, identifying the target acquisition location and detection object, and using a preset segmentation network to identify food categories and remaining information, dietary intake monitoring information is established.

Benefits of technology

It achieves seamless monitoring of food intake, automatically locates the target in the dining scene, improves the efficiency of food intake monitoring, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497035B_ABST
    Figure CN115497035B_ABST
Patent Text Reader

Abstract

The application relates to a diet intake monitoring method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a dining scene image set of a current collection round, determining at least one target collection orientation corresponding to the current collection round according to the dining scene image set, the target collection orientation is used for positioning a region where a main detection object is located, determining at least one target detection object according to a target positioning point corresponding to each target collection orientation, the target positioning point is used for positioning a sub-detection object, obtaining a diet collection array corresponding to the current collection round according to diet recognition information corresponding to each target detection object, the diet recognition information comprises food categories and residual information corresponding to each food category, and obtaining diet intake monitoring information based on the diet collection array corresponding to each collection round in a preset dining time range. The method can automatically position each detection object for diet intake analysis, and the diet intake monitoring efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for monitoring dietary intake. Background Technology

[0002] The health status of children's diet in kindergartens has always been a major concern for parents. Due to the limited number of teachers, it is impossible to take care of all children at the same time, and young children have relatively low self-discipline in eating. Therefore, how to obtain information about children's diet is an urgent problem to be solved.

[0003] Traditional methods typically involve viewing videos or taking pictures for feedback, but these methods cannot obtain specific information about children's food intake during meals. Alternatively, food testing methods that require manual intervention can be used, but these methods increase the children's mealtime process and require teachers to process the data afterward, thus increasing labor costs.

[0004] Therefore, the relevant technologies suffer from low efficiency in monitoring dietary intake. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, apparatus, computer equipment, storage medium, and computer program product for monitoring dietary intake that can solve the above-mentioned technical problems.

[0006] In a first aspect, this application provides a method for monitoring dietary intake, the method comprising:

[0007] Obtain the dining scene image set for the current collection round, and determine at least one target collection direction corresponding to the current collection round based on the dining scene image set; each dining scene image in the dining scene image set has a corresponding collection direction, and the target collection direction is used to locate the area where the main detection object is located;

[0008] Based on the target location points corresponding to the target acquisition locations, at least one target detection object is determined; the target location points are used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship;

[0009] Based on the dietary identification information corresponding to each of the target detection objects, a dietary collection array corresponding to the current collection round is obtained; the dietary identification information includes food categories and remaining quantity information corresponding to each food category;

[0010] Based on the dietary data collection array corresponding to each collection round within the preset meal time range, dietary intake monitoring information is obtained.

[0011] In one embodiment, acquiring the dining scene image set for the current acquisition round, and determining at least one target acquisition location corresponding to the current acquisition round based on the dining scene image set, includes:

[0012] The dining scene images captured in real time by the PTZ camera device from different acquisition positions are obtained to obtain the dining scene image set of the current acquisition round;

[0013] The at least one target acquisition location is obtained based on the candidate detection regions identified from each of the dining scene images; the candidate detection regions are the regions in the image where the main detection object is identified.

[0014] In one embodiment, obtaining the at least one target acquisition location based on candidate detection regions identified from each of the dining scene images includes:

[0015] For each candidate detection region, the image coordinate points corresponding to the candidate detection region are converted into three-dimensional coordinate points based on the gimbal camera device;

[0016] Based on the three-dimensional coordinate points corresponding to each candidate detection region, a set of candidate orientations is obtained, and duplicate orientations in the set of candidate orientations are removed to obtain the at least one target acquisition orientation; the duplicate orientations are candidate orientations obtained for the same main detection object.

[0017] After obtaining the at least one target acquisition orientation, the method further includes:

[0018] Path planning is performed using at least one target acquisition location to obtain the path planning result corresponding to the current acquisition round.

[0019] In one embodiment, converting the image coordinates corresponding to the candidate detection region into three-dimensional coordinates based on the gimbal camera device includes:

[0020] The image coordinates corresponding to the candidate detection region are converted into rotational physical quantities based on the gimbal camera device.

[0021] The acquisition position corresponding to the dining scene image to which the candidate detection area belongs is taken as the current acquisition position, and the current acquisition position is converted into a three-dimensional coordinate point based on the pan-tilt camera device;

[0022] By combining the rotational physical quantity corresponding to the candidate detection region and the three-dimensional coordinate point corresponding to the current acquisition orientation, the three-dimensional coordinate point corresponding to the candidate detection region is obtained.

[0023] In one embodiment, determining at least one target detection object based on the target positioning point corresponding to each of the target acquisition azimuths includes:

[0024] For each target acquisition location, the target location point is determined based on the sub-detection objects identified from the area where the main detection object is located;

[0025] Based on the target location point, a weighted bipartite graph is established based on the three-dimensional coordinate distance of the sub-detected object;

[0026] The weighted bipartite graph is used to match the sub-detection objects with their labels, and the at least one target detection object is determined based on the matching results.

[0027] In one embodiment, obtaining the diet collection array corresponding to the current collection round based on the diet identification information corresponding to each of the target detection objects includes:

[0028] For each target detection object, an image of the detection object corresponding to the target detection object is obtained; the target detection object is an item object used to hold food, and the item object has multiple partitioned regions;

[0029] The image of the detected object is input into a first preset segmentation network to obtain a first pixel region corresponding to the object; and the image of the detected object is input into a second preset segmentation network to obtain a second pixel region corresponding to each segmentation region.

[0030] Identify the food category corresponding to each of the second pixel regions, and determine the remaining quantity information corresponding to each food category based on the first pixel region and the second pixel region;

[0031] The food category and the remaining quantity information corresponding to each food category are used as the dietary identification information corresponding to the target detection object.

[0032] Secondly, this application also provides a dietary intake monitoring device, the device comprising:

[0033] The acquisition orientation determination module is used to acquire a set of dining scene images for the current acquisition round, and determine at least one target acquisition orientation corresponding to the current acquisition round based on the set of dining scene images; each dining scene image in the set of dining scene images has a corresponding acquisition orientation, and the target acquisition orientation is used to locate the area where the main detection object is located;

[0034] The detection object determination module is used to determine at least one target detection object based on the target positioning point corresponding to each target acquisition location; the target positioning point is used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship;

[0035] The food collection array acquisition module is used to obtain the food collection array corresponding to the current collection round based on the food identification information corresponding to each of the target detection objects; the food identification information includes food categories and the remaining quantity information corresponding to each food category;

[0036] The dietary intake monitoring information acquisition module is used to obtain dietary intake monitoring information based on the dietary collection array corresponding to each collection round within a preset meal time range.

[0037] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the dietary intake monitoring method described above.

[0038] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the dietary intake monitoring method described above.

[0039] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the dietary intake monitoring method described above.

[0040] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for monitoring dietary intake acquire a set of dining scene images for the current collection round. Based on the set of dining scene images, at least one target acquisition location corresponding to the current collection round is determined. Each dining scene image in the set of images has a corresponding acquisition location. The target acquisition location is used to locate the area where the main detection object is located. Then, based on the target location point corresponding to each target acquisition location, at least one target detection object is determined. The target location point is used to locate sub-detection objects. The sub-detection objects exist in the area where the main detection object is located. The target detection objects are sub-detection objects with object association relationships. Furthermore, based on the dietary identification information corresponding to each target detection object, a dietary collection array corresponding to the current collection round is obtained. The dietary identification information includes food categories and the remaining quantity information corresponding to each food category. Based on the dietary collection array corresponding to each collection round within a preset dining time range, dietary intake monitoring information is obtained. This achieves automatic location of each detection object in the dining scene for dietary intake analysis to obtain the ingested food categories and remaining quantity information corresponding to the associated dietary intake objects. It can achieve a completely seamless effect and improve the efficiency of dietary intake monitoring. Attached Figure Description

[0041] Figure 1 This is a flowchart illustrating a dietary intake monitoring method in one embodiment;

[0042] Figure 2a This is a schematic diagram of a dining scene image in one embodiment;

[0043] Figure 2b This is a schematic diagram of a cruise procedure in one embodiment;

[0044] Figure 2c This is a schematic diagram illustrating one method of analyzing detection results in one embodiment;

[0045] Figure 3a This is a schematic diagram of the coordinate system of a gimbal camera device in one embodiment;

[0046] Figure 3b This is a schematic diagram of an image coordinate system in one embodiment;

[0047] Figure 3c This is a schematic diagram of a rotation angle about the Y-axis in one embodiment;

[0048] Figure 4 This is a flowchart illustrating another dietary intake monitoring method in one embodiment;

[0049] Figure 5 This is a structural block diagram of a dietary intake monitoring device in one embodiment;

[0050] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0052] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties; correspondingly, this application also provides a corresponding user authorization entry point for users to choose to authorize or refuse.

[0053] In one embodiment, such as Figure 1 As shown, a method for monitoring dietary intake is provided. This embodiment illustrates the application of this method to a server. It is understood that this method can also be applied to a terminal, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0054] Step 101: Obtain the dining scene image set for the current collection round; determine at least one target collection location corresponding to the current collection round based on the dining scene image set; each dining scene image in the dining scene image set has a corresponding collection location, and the target collection location is used to locate the area where the main detection object is located.

[0055] Specifically, for dining scenarios, multiple rounds of image acquisition tasks can be configured based on a preset dining time range. Then, the pan-tilt camera device can perform multiple rounds of image cyclic acquisition to obtain data from multiple rounds within the preset dining time range. The current acquisition round can correspond to any round of image acquisition tasks.

[0056] As an example, the target acquisition direction can be the movement direction of the gimbal camera. For example, the main cruise point of the gimbal camera during the cruise process can be configured according to the rotation angle of the gimbal camera.

[0057] In practical applications, a set of dining scene images for the current collection round can be acquired using a pan-tilt camera device. Since each dining scene image in the set has a corresponding acquisition orientation, target detection and localization can be performed based on the set of dining scene images. In this way, at least one target acquisition orientation can be determined for further positioning and focusing for identification. This target acquisition orientation can be used to locate the area where the main detection object is located, such as the area where the table is located.

[0058] Specifically, it can acquire real-time images of the dining scene captured by the PTZ camera device from different acquisition positions (such as...). Figure 2a (As shown), then target detection and localization can be performed on each dining scene image to determine candidate detection regions, such as... Figure 2a In the table area, and then for each candidate detection area, by converting the image coordinate points corresponding to the candidate detection area into three-dimensional coordinate points based on the pan-tilt camera device, and using the IOU (Intersection-over-Union) method to eliminate the acquisition orientation of the same main detection object that is repeatedly located, at least one target acquisition orientation can be obtained.

[0059] In one example, such as Figure 2b The illustrated cruise process shows that, for the current data collection round, images of the dining scene can be collected using a gimbal camera. Then, a cruise map can be calculated based on the collected images to obtain the path planning results for the current data collection round. The target location can then be used as the main cruise point for further cruise detection and recognition.

[0060] Step 102: Based on the target positioning points corresponding to the target acquisition locations of each target, determine at least one target detection object; the target positioning points are used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship;

[0061] The target positioning point can be determined based on the sub-detected objects identified by target detection. For example, the center coordinates of each sub-detected object can be converted into the rotation angle of the pan-tilt camera, which can then be used as the sub-cruising point (i.e., the target positioning point) of the pan-tilt camera during its patrol process.

[0062] As an example, by performing target detection and recognition on the area where the main detection object is located, such as detecting a plate and a specific number plate in the table area, further matching can be performed based on the identified sub-detection objects existing in the area where the main detection object is located. For example, the plate can be associated with a specific number plate, and thus the plate associated with the specific number plate can be used as the target detection object. Figure 2a As shown.

[0063] In the specific implementation, for each target collection location, by performing target detection and identification on the area where the main detection object is located, the target location point corresponding to the target collection location can be obtained based on the identified sub-detection objects. Then, the sub-detection objects can be further matched based on the target location point, and the sub-detection objects with object association relationships can be used as target detection objects. Thus, the number plate can be used to bind the plate to each food consumption object.

[0064] In an optional embodiment, for each target acquisition location, the target location point can be determined based on the sub-detection objects identified from the area where the main detection object is located. Then, based on the target location point, a weighted bipartite graph based on the three-dimensional coordinate distance of the sub-detection objects can be established. The weighted bipartite graph can then be used to match the sub-detection objects with their labels. For example, the Hungarian matching algorithm can be used to calculate the matching situation between a plate and a specific number plate, and the target detection object can be determined based on the matching results.

[0065] For example, such as Figure 2b As shown, target detection is performed for each main cruise point (i.e., target acquisition location) to coarsely locate the plate and number plate, and obtain the sub-cruise points (i.e. target location points). Then, by traversing the sub-cruise points, the plate and number plate (i.e. sub-detection objects) are matched to obtain the paired plate and number plate, i.e. target detection objects.

[0066] Step 103: Based on the dietary identification information corresponding to each of the target detection objects, obtain the dietary collection array corresponding to the current collection round; the dietary identification information includes food categories and the remaining quantity information corresponding to each food category;

[0067] The target detection object can be an item object used to hold food, such as a plate. This item object can have multiple divided areas, such as the compartments in the plate.

[0068] After obtaining the target detection object, the food category and the remaining quantity of each food category can be extracted from the target detection object to obtain the corresponding food identification information. Then, based on the food identification information corresponding to each target detection object, the food collection array corresponding to the current collection round can be generated.

[0069] Specifically, for each target detection object, the corresponding detection object image can be obtained, such as the segmented plate image. Then, the detection object image can be input into a first preset segmentation network to obtain the first pixel region corresponding to the object, such as the plate pixel region. The detection object image can also be input into a second preset segmentation network to obtain the second pixel region corresponding to each segmentation region, such as the grid pixel region. Then, the food category corresponding to each second pixel region can be identified, and the remaining information corresponding to each food category, such as the amount of food remaining, can be determined based on the first pixel region and the second pixel region.

[0070] In one example, such as Figure 2cThe analysis and detection results shown can determine the second pixel region corresponding to each segmentation area based on the image of the detected object, such as the pixel region of the grid in a plate. For each second pixel region, the food type and the amount of food remaining in the grid can be obtained. The amount of food remaining can be calculated based on the area of ​​the food.

[0071] Step 104: Based on the dietary data collection array corresponding to each collection round within the preset meal time range, obtain dietary intake monitoring information.

[0072] In practical applications, by traversing the main cruise point and sub-cruise point for each collection round, the food collection array corresponding to each collection round can be obtained. Then, based on the food collection array obtained within the preset dining time range, food intake monitoring information in the dining scenario can be generated, such as the type of food and the amount of food remaining after each food intake object eats during the entire dining time period, as well as the type of food and the amount of food consumed by each food intake object.

[0073] For example, the dietary intake monitoring method in this embodiment can be applied to the dining scene of kindergarten children. By using a pan-tilt zoom camera to collect images of the dining scene in a loop, and by using algorithms such as path planning, Hungarian matching, axis-angle rotation, and deep learning, it can automatically monitor and identify the amount of food consumed by children during their meals, providing parents with an objective and visual observation platform.

[0074] In one optional embodiment, the nutritional components ingested by the dietary recipient can be calculated based on the corresponding nutritional data preset for each food category and the remaining information corresponding to each food category. Alternatively, a dietary data collection array obtained by cyclically collecting data throughout the entire meal time period can be used to generate a food quantity line chart for each plate, which can be used to characterize the changes in food quantity of the dietary recipient associated with the plate during the meal time period.

[0075] Compared to traditional methods that rely solely on camera monitoring to observe children's mealtimes (which cannot capture the food and portion sizes, and suffer from blind spots and pixelation issues at distances), or manual methods that combine tableware capacity with photographic images for food identification, this approach requires cooperation between children and teachers, increasing the mealtime process and labor costs. The technical solution of this embodiment, through autonomous positioning and focusing detection of objects, identifies the food categories on the plates and analyzes the amount of food remaining. For example, it can autonomously locate and focus on each child's plate in the dining scene, identify the food categories and the amount of food remaining on the plates, and establish a relationship between the plates and each child. Thus, after binding food categories and nutritional information, it can calculate the intake of nutrients based on the amount of food remaining. A pan-tilt camera device is used to obtain a larger detection range, and a clearer target image is obtained through zooming, so that the plates can be bound to each child based on the number tag. By cyclically traversing the patrol point positions, the amount of food remaining on all plates within the larger detection range can be calculated periodically, and the target data can be obtained autonomously and imperceptibly within a preset time range. This achieves the effect of not changing the dining process and being imperceptible throughout, reducing labor costs, improving the overall efficiency of the food intake analysis process, and also being able to promptly detect problems that occur during the dining process based on the monitoring situation.

[0076] In the aforementioned dietary intake monitoring method, a set of dining scene images for the current collection round is acquired. Based on this set, at least one target acquisition location corresponding to the current collection round is determined. Each dining scene image in the set has a corresponding acquisition location, and the target acquisition location is used to locate the area where the main detection object is located. Then, based on the target location point corresponding to each target acquisition location, at least one target detection object is determined. The target location point is used to locate sub-detection objects, which exist in the area where the main detection object is located. The target detection objects are sub-detection objects with object association relationships. Furthermore, based on the dietary identification information corresponding to each target detection object, a dietary collection array corresponding to the current collection round is obtained. The dietary identification information includes food categories and the remaining quantity information corresponding to each food category. Based on the dietary collection array corresponding to each collection round within a preset dining time range, dietary intake monitoring information is obtained. This achieves automatic location of each detection object in the dining scene for dietary intake analysis to obtain the ingested food categories and remaining quantity information corresponding to the associated dietary intake objects. This achieves a completely seamless effect and improves the efficiency of dietary intake monitoring.

[0077] In one embodiment, acquiring the set of dining scene images for the current acquisition round, and determining at least one target acquisition location corresponding to the current acquisition round based on the set of dining scene images, may include the following steps:

[0078] The dining scene images are acquired in real time from different acquisition positions by the pan-tilt camera device to obtain the dining scene image set of the current acquisition round; the at least one target acquisition position is obtained based on the candidate detection regions identified from each of the dining scene images; the candidate detection region is the region in the image where the main detection object is identified.

[0079] In practical applications, a gimbal camera (such as a gimbal zoom camera) can be positioned in the center of the dining scene. By rotating the gimbal camera to traverse the pitch and yaw angles, images of the dining scene can be acquired in real time from different acquisition positions, resulting in a set of dining scene images for the current acquisition round. Then, candidate detection regions can be extracted based on the dining scene image set. For example, by using object detection to locate the area where the table is located in all images, the ROI (region of interest) in the image can be obtained, that is, the region in the image where the main detection object is identified (such as the table area). Then, for each candidate detection region, the image coordinate points corresponding to the candidate detection region are converted into three-dimensional coordinate points based on the gimbal camera, and the Intersection over Union (IOU) method is used to eliminate acquisition positions that repeatedly locate the same main detection object, thus obtaining at least one target acquisition position.

[0080] In this embodiment, by acquiring dining scene images captured in real time from different acquisition positions by a gimbal camera device, a set of dining scene images for the current acquisition round is obtained. Then, based on the candidate detection areas identified from each dining scene image, at least one target acquisition position is obtained, which helps to further locate, focus, and identify the target detection object.

[0081] In one embodiment, obtaining the at least one target acquisition location based on candidate detection regions identified from each of the dining scene images may include the following steps:

[0082] For each candidate detection region, the image coordinate points corresponding to the candidate detection region are converted into three-dimensional coordinate points based on the pan-tilt camera device; a candidate orientation set is obtained based on the three-dimensional coordinate points corresponding to each candidate detection region, and duplicate orientations in the candidate orientation set are removed to obtain the at least one target acquisition orientation; the duplicate orientations are candidate orientations obtained for the same main detection object.

[0083] In one example, by converting the image coordinates corresponding to each candidate detection region into three-dimensional coordinates based on the PTZ camera, a set of candidate orientations can be obtained based on the three-dimensional coordinates corresponding to each candidate detection region. Then, the IOU method can be used to eliminate the acquisition orientations that repeatedly locate the same main detection object (i.e., duplicate orientations) to obtain at least one target acquisition orientation. For example, the IOU method can be used to eliminate the movement orientations of the PTZ camera that repeatedly locate the same table, and the remaining rotation angles of the PTZ camera can be saved as the main cruise points.

[0084] Specifically, for the Region of Interest (ROI) of the table area in the image, its corresponding image coordinates can be converted into three-dimensional coordinates on the spherical surface of the pan-tilt camera's motion unit. That is, based on the three-dimensional coordinates of the pan-tilt camera device, the method of converting image coordinates into three-dimensional coordinates on the camera's spherical surface can transform the ROI obtained by target detection, thus ensuring the uniqueness of the detected target in space.

[0085] After obtaining the at least one target acquisition orientation, the method further includes:

[0086] Path planning is performed using at least one target acquisition location to obtain the path planning result corresponding to the current acquisition round.

[0087] In another example, after removing the movement orientation of the pan-tilt camera that repeatedly positions itself on the same table and saving the remaining pan-tilt camera rotation angles as the main cruise points, the cruise points can be adjusted to add or remove redundant or missing main cruise points (i.e., target acquisition orientations). Then, the shortest path planning of the pan-tilt camera's cruise point path can be performed based on a genetic algorithm, thereby reducing the time spent by the pan-tilt camera rotating and cruised for one cycle and increasing the amount of data acquired on the detected object.

[0088] In this embodiment, for each candidate detection area, the image coordinate points corresponding to the candidate detection area are converted into three-dimensional coordinate points based on the gimbal camera device. Then, based on the three-dimensional coordinate points corresponding to each candidate detection area, a candidate orientation set is obtained, and duplicate orientations in the candidate orientation set are removed to obtain at least one target acquisition orientation. This can further reduce the time spent on the cruise cycle based on the shortest path planning, providing data support for subsequent detection.

[0089] In one embodiment, converting the image coordinate points corresponding to the candidate detection region into three-dimensional coordinate points based on the gimbal camera device may include the following steps:

[0090] The image coordinates corresponding to the candidate detection region are converted into rotational physical quantities based on the gimbal camera device; the acquisition orientation corresponding to the dining scene image to which the candidate detection region belongs is taken as the current acquisition orientation, and the current acquisition orientation is converted into three-dimensional coordinates based on the gimbal camera device; the three-dimensional coordinates corresponding to the candidate detection region are obtained by combining the rotational physical quantities corresponding to the candidate detection region and the three-dimensional coordinates corresponding to the current acquisition orientation.

[0091] In practical applications, converting the image coordinates corresponding to the candidate detection region into three-dimensional coordinates based on the gimbal camera device, such as three-dimensional coordinates on the sphere of the gimbal camera's motion unit, can be done in the following way:

[0092] 1. Convert the image pixel coordinates into the physical quantity of the gimbal camera's rotation;

[0093] Because the rotation of a gimbal camera has two degrees of freedom, one is rotation about an axis parallel to the mounting plane, i.e., rotation about the X-axis; the other is rotation about an axis perpendicular to the mounting plane, i.e., rotation about the Y-axis, such as... Figure 3a As shown, the axis parallel to the mounting plane can be defined as the camera's x-axis, and the axis perpendicular to the x-axis and the direction directly opposite the camera can be defined as the camera's y-axis. Using a right-handed coordinate system as an example, the z-axis orientation can be obtained. The z-axis does not have rotational property, while the y-axis only has theoretical rotation; in practical applications, the gimbal camera rotates around the Y-axis. Since the x and x-axis rotate synchronously each time the Y-axis rotates, they can be considered as the same axis.

[0094] like Figure 3b As shown, let the current orientation of the gimbal camera be rotated by an angle B around the X-axis and by an angle A around the Y-axis, where the unit vector of the x-axis is (1, 0, 0), the unit vector of the y-axis is (0, cosα, sinα), and the unit vector of the optical center orientation c of the gimbal camera is (0, sinα, -cosα). Then, the coordinate system of the image captured by the gimbal camera at the current orientation is parallel to the xOy plane. If the image has no distortion, the horizontal and vertical field of view angles can be F respectively. x F y Then, the coordinates p with the image center o as the origin correspond to the rotation angles α and β around the y-axis and around the x-axis, respectively:

[0095]

[0096] The unit vector c can be rotated by β around the x-axis and then by α around the y-axis, as shown in the following formula:

[0097] c'=ccosβ+(1-cosβ)(x·c)x+(x×c)sinβ,

[0098] c'=c'cosα+(1-cosα)(y·c')y+(y×c')sinα

[0099] Since the x-axis is coaxial with the x-axis, the angle of rotation around the x-axis can be directly obtained as follows:

[0100] Δβ=arcsinc' y -B

[0101] The vector of the gimbal camera after rotation around the x-axis is like Figure 3c As shown, △α in the plane parallel to XOZ is the rotation angle about the Y-axis:

[0102]

[0103] 2. Convert the orientation of the PTZ camera to the three-dimensional coordinates of the PTZ camera's unit sphere;

[0104] Let the current position of the gimbal camera be rotated by an angle B around the X-axis and an angle A around the Y-axis, with the x-axis being (1, 0, 0), the y-axis being (0, 1, 0), and the optical center direction c being (0, 0, -1). Then, the conversion of the three-dimensional coordinates of the current position of the gimbal camera on the unit sphere can be expressed as c rotating by B around the x-axis and then rotating by A around the y-axis.

[0105] 3. Convert the image pixel coordinates to the three-dimensional coordinates of the unit sphere of the gimbal camera;

[0106] By combining the two methods above, the physical quantity of the PTZ camera to be rotated obtained by method 1 is superimposed on the current orientation of the PTZ camera and substituted into method 2 to obtain the target result, namely the three-dimensional coordinate point corresponding to the candidate detection area.

[0107] In this embodiment, by converting the image coordinates corresponding to the candidate detection area into rotational physical quantities based on the gimbal camera, and then taking the acquisition orientation of the dining scene image to which the candidate detection area belongs as the current acquisition orientation, and converting the current acquisition orientation into three-dimensional coordinates based on the gimbal camera, and then combining the rotational physical quantities corresponding to the candidate detection area and the three-dimensional coordinates corresponding to the current acquisition orientation, the three-dimensional coordinates corresponding to the candidate detection area can be obtained, which can ensure the uniqueness of the detected target in space.

[0108] In one embodiment, determining at least one target detection object based on the target positioning point corresponding to each of the target acquisition azimuths may include the following steps:

[0109] For each target acquisition location, the target location point is determined based on the sub-detection objects identified from the area where the main detection object is located; a weighted bipartite graph based on the three-dimensional coordinate distance of the sub-detection objects is established based on the target location point; the weighted bipartite graph is used to match the sub-detection objects with the sub-detection object labels, and the at least one target detection object is determined based on the matching result.

[0110] In one example, when further detecting plates and number plates, the movement orientation corresponding to the area where plates and number plates are detected can be determined based on the main cruise point, i.e., the target acquisition orientation. The plates and specific number plates (i.e., sub-detection objects) existing in the table area can be detected through further target localization. For example, the plates and specific number plates can be configured to meet the positioning conditions and be inside the table ROI. Then, after detecting the plates and number plates, the coordinates of each target center point can be converted into the rotation angle of the pan-tilt camera device as a sub-cruise point, which can be used to further locate the plates and number plates through target detection. That is, the target positioning point is determined based on the sub-detection objects identified from the area where the main detection object is located.

[0111] In another example, during the process of binding a plate to a number plate, after converting the coordinates of the number plate and the plate into spherical 3D coordinates of the pan-tilt camera, duplicate targets are eliminated using the Intersection over Union (IOU) method. Then, a weighted bipartite graph of the 3D coordinate distance between the number plate and the plate can be established. Subsequently, the minimum distance can be found using the Hungarian algorithm to match and bind the plate and number plate. Specifically, it can be determined whether the distance between the number plate and the plate exceeds a reference threshold. If the distance exceeds the reference threshold, it is marked as an invalid match. The threshold can then be applied to the next pair of number plates and plates, thus obtaining at least one plate with a matched number plate (i.e., the target detection object). Therefore, by using the Hungarian matching algorithm to calculate the optimal coordinate matching of the number plate and the plate in spherical space, a unique binding association can be established between the plate with the matched number plate and the object of food consumption.

[0112] In this embodiment, by collecting data from each target location, the target location point is determined based on the sub-detection objects identified from the area where the main detection object is located. Then, based on the target location point, a weighted bipartite graph based on the three-dimensional coordinate distance of the sub-detection objects is established. The weighted bipartite graph is then used to match the sub-detection objects with their labels, and at least one target detection object is determined based on the matching results. This enables effective matching of number plates and plates, achieving a unique binding association between the plate and the food consumption object after matching, thus improving the accuracy of detection and analysis.

[0113] In one embodiment, obtaining the diet collection array corresponding to the current collection round based on the diet identification information corresponding to each of the target detection objects may include the following steps:

[0114] For each target detection object, an image of the target detection object is acquired; the target detection object is an object used to hold food, and the object has multiple segmented regions; the image of the detection object is input into a first preset segmentation network to obtain a first pixel region corresponding to the object, and the image of the detection object is input into a second preset segmentation network to obtain a second pixel region corresponding to each segmented region; the food category corresponding to each second pixel region is identified, and the remaining quantity information corresponding to each food category is determined based on the first pixel region and the second pixel region; the food category and the remaining quantity information corresponding to each food category are used as the dietary recognition information corresponding to the target detection object.

[0115] In practical applications, the amount and category of food remaining can be extracted for each target detection object (such as a plate). The plate image (i.e., the detected object image) can be obtained by segmenting the ROI of the plate using object detection. Then, the plate image can be input into a semantic segmentation network (i.e., the first preset segmentation network) and an instance segmentation network (i.e., the second preset segmentation network). The semantic segmentation network can be used to label the pixel region Q (i.e., the first pixel region) belonging to the plate, and the instance segmentation network can be used to extract the pixel region P (i.e., the second pixel region) of each compartment. Then, the compartment region sub-image can be input into a metric learning network to obtain the identified food category. The food remaining amount mask can be obtained through P&~Q operation, and the pixel equivalent (i.e., the remaining amount information corresponding to each food category) can be obtained according to the side length and depth of the plate to calculate the amount of food remaining.

[0116] In one example, by using semantic segmentation to extract plate pixels and instance segmentation to extract grid pixels, and using grid regions to infer food categories, and using the common areas of non-plate pixels and grid pixels to extract the amount of food remaining on the plate, an automated process can be achieved from panoramic dining images captured by a PTZ camera to extract the food category and remaining amount of each plate.

[0117] In this embodiment, for each target detection object, an image of the target detection object is acquired, and then the image is input into a first preset segmentation network to obtain a first pixel region corresponding to the object. The image is also input into a second preset segmentation network to obtain a second pixel region corresponding to each segmentation region. The food category corresponding to each second pixel region is identified. Based on the first and second pixel regions, the remaining quantity information corresponding to each food category is determined. The food category and the remaining quantity information corresponding to each food category are then used as the dietary identification information corresponding to the target detection object. This allows for automatic location of each detection object in a dining scenario to perform dietary intake analysis, thereby improving the efficiency of dietary intake monitoring.

[0118] In one embodiment, such as Figure 4 The diagram illustrates another method for monitoring dietary intake. In this embodiment, the method includes the following steps:

[0119] In step 401, dining scene images captured in real time by the pan-tilt camera from different acquisition positions are obtained to form the dining scene image set for the current acquisition round. In step 402, at least one target acquisition position is obtained based on candidate detection regions identified from each dining scene image; the candidate detection region is the region in the image where the main detection object is identified. In step 403, path planning is performed using at least one target acquisition position to obtain the path planning result corresponding to the current acquisition round. In step 404, at least one target detection object is determined based on the target positioning point corresponding to each target acquisition position; the target positioning point is used to locate the sub-detection object. In step 405, for each target detection object, the corresponding detection object image is obtained; the target detection object is an item object used to hold food, and the item object has multiple segmented regions. In step 406, the detection object image is input into a first preset segmentation network to obtain the first pixel region corresponding to the item object, and the detection object image is input into a second preset segmentation network to obtain the second pixel region corresponding to each segmented region. In step 407, the food category corresponding to each second pixel region is identified, and the remaining quantity information corresponding to each food category is determined based on the first pixel region and the second pixel region. In step 408, the food category and the remaining quantity information corresponding to each food category are used as the dietary identification information corresponding to the target detection object. In step 409, dietary intake monitoring information is obtained based on the dietary collection array corresponding to each collection round within a preset meal time range. It should be noted that the specific limitations of the above steps can be found in the specific limitations of a dietary intake monitoring method described above, and will not be repeated here.

[0120] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0121] Based on the same inventive concept, this application also provides a diet intake monitoring device for implementing the diet intake monitoring method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more diet intake monitoring device embodiments provided below can be found in the limitations of the diet intake monitoring method described above, and will not be repeated here.

[0122] In one embodiment, such as Figure 5 As shown, a diet intake monitoring device is provided, comprising:

[0123] The acquisition orientation determination module 501 is used to acquire a set of dining scene images for the current acquisition round, and determine at least one target acquisition orientation corresponding to the current acquisition round based on the set of dining scene images; each dining scene image in the set of dining scene images has a corresponding acquisition orientation, and the target acquisition orientation is used to locate the area where the main detection object is located;

[0124] The detection object determination module 502 is used to determine at least one target detection object based on the target positioning point corresponding to each target acquisition location; the target positioning point is used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship;

[0125] The food collection array acquisition module 503 is used to obtain the food collection array corresponding to the current collection round based on the food identification information corresponding to each of the target detection objects; the food identification information includes food categories and the remaining quantity information corresponding to each food category;

[0126] The dietary intake monitoring information acquisition module 504 is used to obtain dietary intake monitoring information based on the dietary collection array corresponding to each collection round within a preset meal time range.

[0127] In one embodiment, the acquisition orientation determination module 501 includes:

[0128] The image acquisition submodule is used to acquire dining scene images captured in real time by the PTZ camera device from different acquisition positions, and obtain the dining scene image set for the current acquisition round.

[0129] The acquisition orientation submodule is used to obtain the at least one target acquisition orientation based on the candidate detection regions identified from each of the dining scene images; the candidate detection regions are the regions in the image where the main detection object is identified.

[0130] In one embodiment, the location acquisition submodule includes:

[0131] The coordinate point conversion unit is used to convert the image coordinate points corresponding to each candidate detection area into three-dimensional coordinate points based on the pan-tilt camera device for each candidate detection area.

[0132] The duplicate orientation removal unit is used to obtain a set of candidate orientations based on the three-dimensional coordinate points corresponding to each of the candidate detection regions, and to remove duplicate orientations from the set of candidate orientations to obtain the at least one target acquisition orientation; the duplicate orientations are candidate orientations obtained for the same main detection object;

[0133] Also includes:

[0134] The path planning module is used to perform path planning using the at least one target acquisition orientation to obtain the path planning result corresponding to the current acquisition round.

[0135] In one embodiment, the coordinate point transformation unit includes:

[0136] The physical quantity conversion subunit is used to convert the image coordinate points corresponding to the candidate detection area into rotational physical quantities based on the gimbal camera device.

[0137] The acquisition orientation conversion subunit is used to take the acquisition orientation corresponding to the dining scene image to which the candidate detection area belongs as the current acquisition orientation, and convert the current acquisition orientation into a three-dimensional coordinate point based on the pan-tilt camera device;

[0138] The three-dimensional coordinate point is obtained as a sub-unit, which is used to combine the rotational physical quantity corresponding to the candidate detection area and the three-dimensional coordinate point corresponding to the current acquisition orientation to obtain the three-dimensional coordinate point corresponding to the candidate detection area.

[0139] In one embodiment, the detection object determination module 502 includes:

[0140] The target location point determination submodule is used to determine the target location point for each target acquisition orientation based on the sub-detection objects identified from the area where the main detection object is located;

[0141] The matching graph establishment submodule is used to establish a weighted bipartite graph based on the three-dimensional coordinate distance of the sub-detected object according to the target positioning point;

[0142] The matching submodule is used to match the sub-detection object with the sub-detection object label using the weighted bipartite graph, and to determine the at least one target detection object based on the matching result.

[0143] In one embodiment, the diet collection array obtaining module 503 includes:

[0144] The detection object image acquisition submodule is used to acquire the detection object image corresponding to each target detection object; the target detection object is an item object used to hold food, and the item object has multiple partitioned regions;

[0145] The segmentation network processing submodule is used to input the detected object image into a first preset segmentation network to obtain a first pixel region corresponding to the object, and to input the detected object image into a second preset segmentation network to obtain a second pixel region corresponding to each segmentation region.

[0146] The identification and analysis submodule is used to identify the food category corresponding to each of the second pixel regions, and determine the remaining quantity information corresponding to each food category based on the first pixel region and the second pixel region.

[0147] The food identification information acquisition submodule is used to take the food category and the remaining quantity information corresponding to each food category as the food identification information corresponding to the target detection object.

[0148] Each module in the aforementioned dietary intake monitoring device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0149] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores dietary intake monitoring data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a dietary intake monitoring method.

[0150] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0151] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0152] Obtain the dining scene image set for the current collection round, and determine at least one target collection direction corresponding to the current collection round based on the dining scene image set; each dining scene image in the dining scene image set has a corresponding collection direction, and the target collection direction is used to locate the area where the main detection object is located;

[0153] Based on the target location points corresponding to the target acquisition locations, at least one target detection object is determined; the target location points are used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship;

[0154] Based on the dietary identification information corresponding to each of the target detection objects, a dietary collection array corresponding to the current collection round is obtained; the dietary identification information includes food categories and remaining quantity information corresponding to each food category;

[0155] Based on the dietary data collection array corresponding to each collection round within the preset meal time range, dietary intake monitoring information is obtained.

[0156] In one embodiment, the processor, when executing a computer program, also implements the steps of the dietary intake monitoring method in the other embodiments described above.

[0157] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0158] Obtain the dining scene image set for the current collection round, and determine at least one target collection direction corresponding to the current collection round based on the dining scene image set; each dining scene image in the dining scene image set has a corresponding collection direction, and the target collection direction is used to locate the area where the main detection object is located;

[0159] Based on the target location points corresponding to the target acquisition locations, at least one target detection object is determined; the target location points are used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship;

[0160] Based on the dietary identification information corresponding to each of the target detection objects, a dietary collection array corresponding to the current collection round is obtained; the dietary identification information includes food categories and remaining quantity information corresponding to each food category;

[0161] Based on the dietary data collection array corresponding to each collection round within the preset meal time range, dietary intake monitoring information is obtained.

[0162] In one embodiment, when the computer program is executed by a processor, it also implements the steps of the dietary intake monitoring method in the other embodiments described above.

[0163] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0164] Obtain the dining scene image set for the current collection round, and determine at least one target collection direction corresponding to the current collection round based on the dining scene image set; each dining scene image in the dining scene image set has a corresponding collection direction, and the target collection direction is used to locate the area where the main detection object is located;

[0165] Based on the target location points corresponding to the target acquisition locations, at least one target detection object is determined; the target location points are used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship;

[0166] Based on the dietary identification information corresponding to each of the target detection objects, a dietary collection array corresponding to the current collection round is obtained; the dietary identification information includes food categories and remaining quantity information corresponding to each food category;

[0167] Based on the dietary data collection array corresponding to each collection round within the preset meal time range, dietary intake monitoring information is obtained.

[0168] In one embodiment, when the computer program is executed by a processor, it also implements the steps of the dietary intake monitoring method in the other embodiments described above.

[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0170] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0171] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for monitoring dietary intake, characterized in that, The method includes: The process involves: acquiring a set of dining scene images for the current acquisition round; determining at least one target acquisition location corresponding to the current acquisition round based on the set of dining scene images; each dining scene image in the set of dining scene images has a corresponding acquisition location, and the target acquisition location is used to locate the area where the main detection object is located; including: acquiring dining scene images acquired in real time by a pan-tilt camera device from different acquisition locations to obtain the set of dining scene images for the current acquisition round; obtaining the at least one target acquisition location based on candidate detection regions identified from each of the dining scene images; the candidate detection region is the region in the image where the main detection object is identified; Based on the target location points corresponding to the target acquisition locations, at least one target detection object is determined; the target location points are used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship; Based on the dietary recognition information corresponding to each of the target detection objects, a dietary collection array corresponding to the current collection round is obtained; the dietary recognition information includes food categories and remaining quantity information corresponding to each food category; including: for each target detection object, acquiring the detection object image corresponding to the target detection object; the target detection object is an object used to hold food, and the object has multiple segmented regions; inputting the detection object image into a first preset segmentation network to obtain the pixel region corresponding to the object, as the first pixel region, and inputting the detection object image into a second preset segmentation network to obtain the pixel region corresponding to each segmented region, as each second pixel region; identifying each second pixel. The system identifies the food category corresponding to the region and determines the remaining quantity information corresponding to each food category based on the first pixel region and the second pixel region. The food category and the remaining quantity information corresponding to each food category are used as the dietary recognition information corresponding to the target detection object. The first preset segmentation network is a semantic segmentation network; the second preset segmentation network is an instance segmentation network. The system includes: inputting the second pixel region into a metric learning network to obtain the identified food category; obtaining a food remaining quantity mask through P&Q operations to obtain pixel equivalents based on the side length and depth of the target detection object, which are used as the remaining quantity information corresponding to each food category; where P is the second pixel region and Q is the first pixel region. Based on the dietary data collection array corresponding to each collection round within the preset meal time range, dietary intake monitoring information is obtained.

2. The method according to claim 1, characterized in that, The step of obtaining the at least one target acquisition location based on candidate detection regions identified from each of the dining scene images includes: For each candidate detection region, the image coordinate points corresponding to the candidate detection region are converted into three-dimensional coordinate points based on the gimbal camera device; Based on the three-dimensional coordinate points corresponding to each candidate detection region, a set of candidate orientations is obtained, and duplicate orientations in the set of candidate orientations are removed to obtain the at least one target acquisition orientation; the duplicate orientations are candidate orientations obtained for the same main detection object. After obtaining the at least one target acquisition orientation, the method further includes: Path planning is performed using at least one target acquisition location to obtain the path planning result corresponding to the current acquisition round.

3. The method according to claim 2, characterized in that, The step of converting the image coordinate points corresponding to the candidate detection region into three-dimensional coordinate points based on the gimbal camera device includes: The image coordinates corresponding to the candidate detection region are converted into rotational physical quantities based on the gimbal camera device. The acquisition position corresponding to the dining scene image to which the candidate detection area belongs is taken as the current acquisition position, and the current acquisition position is converted into a three-dimensional coordinate point based on the pan-tilt camera device; By combining the rotational physical quantity corresponding to the candidate detection region and the three-dimensional coordinate point corresponding to the current acquisition orientation, the three-dimensional coordinate point corresponding to the candidate detection region is obtained.

4. The method according to claim 1, characterized in that, The step of determining at least one target detection object based on the target positioning point corresponding to each of the target acquisition azimuths includes: For each target acquisition location, the target location point is determined based on the sub-detection objects identified from the area where the main detection object is located; Based on the target location point, a weighted bipartite graph is established based on the three-dimensional coordinate distance of the sub-detected object; The weighted bipartite graph is used to match the sub-detection objects with their labels, and the at least one target detection object is determined based on the matching results.

5. A dietary intake monitoring device, characterized in that, The device includes: The acquisition orientation determination module is used to acquire a set of dining scene images for the current acquisition round, and determine at least one target acquisition orientation corresponding to the current acquisition round based on the set of dining scene images; each dining scene image in the set of dining scene images has a corresponding acquisition orientation, and the target acquisition orientation is used to locate the area where the main detection object is located; The acquisition orientation determination module is specifically used to acquire dining scene images captured in real time by the pan-tilt camera device from different acquisition orientations, and obtain the dining scene image set of the current acquisition round; based on the candidate detection regions identified from each of the dining scene images, the at least one target acquisition orientation is obtained; the candidate detection region is the region in the image where the main detection object is identified; The detection object determination module is used to determine at least one target detection object based on the target positioning point corresponding to each target acquisition location; the target positioning point is used to locate sub-detection objects, the sub-detection objects exist in the area where the main detection object is located, and the target detection object is a sub-detection object with object association relationship; The food collection array acquisition module is used to obtain the food collection array corresponding to the current collection round based on the food identification information corresponding to each of the target detection objects; the food identification information includes food categories and the remaining quantity information corresponding to each food category; The food collection array acquisition module is specifically used to acquire the detection object image corresponding to each target detection object; the target detection object is an item object used to hold food, and the item object has multiple segmented regions; the detection object image is input into a first preset segmentation network to obtain the pixel region corresponding to the item object, as the first pixel region; and the detection object image is input into a second preset segmentation network to obtain the pixel region corresponding to each segmented region, as each second pixel region; the food category corresponding to each second pixel region is identified, and the remaining quantity information corresponding to each food category is determined based on the first pixel region and the second pixel region; the food category and the remaining quantity information corresponding to each food category are used as the food recognition information corresponding to the target detection object; the first preset segmentation network is a semantic segmentation network; the second preset segmentation network is an instance segmentation network; The food collection array acquisition module is specifically used to input the second pixel region into the metric learning network to obtain the identified food category; obtain the food remaining amount mask through P&Q operation, and obtain the pixel equivalent according to the side length and depth of the target detection object, as the remaining amount information corresponding to each food category; wherein, P is the second pixel region and Q is the first pixel region. The dietary intake monitoring information acquisition module is used to obtain dietary intake monitoring information based on the dietary collection array corresponding to each collection round within a preset meal time range.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.