Illumination control method and device, computer equipment and storage medium

By analyzing the user's behavior patterns and using image processing and action recognition technology to determine lighting needs, the problem that existing lighting control technologies fail to meet the needs of different behavior patterns is solved, and a higher degree of intelligence and convenience is achieved.

CN119997316APending Publication Date: 2025-05-13ARCSOFT CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411891673.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing lighting control technology fails to effectively consider the lighting needs of users under different behavioral modes, resulting in a low degree of intelligence.

Method used

By acquiring the preambled continuous multi-frame image of the object to be detected, its target behavior state is determined, including the first target behavior state and the second target behavior state. The first target behavior state is determined by hand information and posture information, and the second target behavior state is determined by action recognition model. The lighting state of the lighting device is controlled according to the target behavior state.

Benefits of technology

It realizes accurate judgment of lighting needs based on user behavior patterns, improves the intelligence of lighting control, and provides a more convenient and comfortable lighting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119997316A_ABST
    Figure CN119997316A_ABST
Patent Text Reader

Abstract

The invention relates to a lighting control method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a plurality of preorder continuous frames of images of an object to be detected; according to the preorder continuous multi-frame images, target behavior states of the to-be-detected object are determined, and the target behavior states comprise a first target behavior state and / or a second target behavior state; wherein the first target behavior state is determined by acquiring hand information and posture information of the to-be-detected object in a to-be-illuminated area according to the preorder continuous multi-frame image, and determining the first target behavior state according to the hand information and the posture information; the second target behavior state is determined according to the preorder continuous multi-frame image; and controlling the lighting state of a lighting device according to the target behavior state. According to the invention, the intelligent degree of illumination control is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of lighting control, and in particular to a lighting control method, device, computer equipment and storage medium. Background Art

[0002] Lighting devices have been widely used in homes, public places, inside and outside vehicles and other areas, but the control method of lighting devices is usually more traditional, usually need to be manually turned on or off. If the ideas and methods of intelligent control can be incorporated into the lighting control process, it can provide users with a more convenient and comfortable lighting experience.

[0003] In traditional technology, lighting control is usually automatically adjusted according to the brightness of the area to be illuminated and the user-defined lighting time, without considering the lighting needs of users under different behavior modes, and the degree of intelligence of lighting control is low. Summary of the invention

[0004] Based on this, it is necessary to provide a lighting control method, device, computer equipment and storage medium that can improve the intelligence level of lighting control in response to the above technical problems.

[0005] In a first aspect, the present application provides a lighting control method. The method comprises:

[0006] Acquire a plurality of consecutive frames of images of the object to be detected;

[0007] Determining a target behavior state of the object to be detected according to the preceding continuous multiple frames of images, wherein the target behavior state includes a first target behavior state and / or a second target behavior state;

[0008] Wherein, the first target behavior state is determined by obtaining the hand information and posture information of the object to be detected in the area to be illuminated according to the preceding continuous multiple frames of images, and according to the hand information and the posture information;

[0009] The second target behavior state is determined according to the preceding continuous multiple frames of images;

[0010] According to the target behavior state, the lighting state of the lighting device is controlled.

[0011] In one embodiment, the method further comprises:

[0012] One or more sub-areas to be illuminated are determined in the area to be illuminated, wherein the sub-areas to be illuminated include at least one of the following: a preset area for objects to be picked up and an area for reading objects.

[0013] In one of the embodiments, the hand information includes at least one of the following: position information and movement information of one or more hands, and the posture information includes at least one of the following: line of sight information, head posture information, and human body posture information.

[0014] In one embodiment, the first target behavior state is determined according to the following method:

[0015] According to the hand information, determining whether there is a hand in the sub-area to be illuminated;

[0016] When there is no hand in the sub-area to be illuminated, the first target behavior state is a state where no illumination is required;

[0017] When there is a hand in the sub-area to be illuminated, a first target behavior state of the object to be detected is determined according to the posture information.

[0018] In one embodiment, determining the first target behavior state of the object to be detected according to the posture information includes:

[0019] Determining a sight direction according to the sight information;

[0020] Determining the head direction according to the head posture information; and / or

[0021] Determining a body direction according to the human body posture information;

[0022] The first target behavior state is determined according to the sight direction, the head direction and / or the body direction.

[0023] In one embodiment, determining the first target behavior state according to the sight direction, the head direction and / or the body direction includes:

[0024] When at least one of the sight line direction, the head direction and the body direction of the object to be detected points to the sub-region to be illuminated where the hand is located, it is determined that the first target behavior state is a state requiring illumination.

[0025] In one embodiment, determining the first target behavior state according to the sight direction, the head direction and / or the body inclination direction further includes:

[0026] When the sight direction, the head direction and the body direction of the object to be detected do not point to the sub-area to be illuminated where the hand is located, it is determined that the first target behavior state is a state where no illumination is required.

[0027] In one of the embodiments, when the sub-area to be illuminated where the hand is located is the reading area, determining the first target behavior state according to the sight direction, the head direction and / or the body direction further includes:

[0028] When the sight line direction and the head direction of the object to be detected do not point to the sub-area to be illuminated where the hand is located, and the duration is not less than the first time threshold, it is determined that the first target behavior state is a state that does not require illumination.

[0029] In one embodiment, the second target behavior state is determined according to the following method:

[0030] Input the preceding continuous multi-frame images into an action recognition model, and output a target recognition action of the object to be detected, wherein the action recognition model includes a text encoder, a video encoder, and a multi-task decoder, wherein the text encoder is used to extract text features of a preset action recognition text, the video encoder is used to extract image features of the preceding continuous multi-frame images, and the multi-task decoder is used to determine the target recognition action according to the text features and the image features;

[0031] The second target behavior state is determined according to the target recognition action.

[0032] In one embodiment, the determining the target recognition action according to the text feature and the image feature includes:

[0033] Classifying the text features of the preset action recognition text to determine the action category label;

[0034] The action category label is matched with the image feature based on contrastive learning, and the target recognition action is determined according to the matching result.

[0035] In one embodiment, the second target behavior state is further determined according to the following method:

[0036] The preceding multiple consecutive frames of images are input into an action recognition sub-model, and a target recognition action of the object to be detected is outputted. The action recognition sub-model is trained based on the text features and the image features.

[0037] In one embodiment, controlling the lighting state of the lighting device according to the target behavior state includes:

[0038] According to the target behavior state, controlling the lighting state of the sub-lighting device corresponding to the sub-area to be illuminated;

[0039] When the target behavior state is a lighting-required state, controlling the sub-lighting device corresponding to the sub-area to be illuminated to be in an on state;

[0040] When the target behavior state is a state where lighting is not required, the sub-lighting device corresponding to the sub-area to be illuminated is controlled to be in a closed state.

[0041] In a second aspect, the present application also provides a lighting control device. The device comprises:

[0042] An image acquisition module is used to acquire a plurality of consecutive frames of images of the object to be detected;

[0043] A target behavior state determination module, used to determine the target behavior state of the object to be detected according to the preceding continuous multiple frames of images, wherein the target behavior state includes a first target behavior state and / or a second target behavior state; wherein the first target behavior state is determined by obtaining hand information and posture information of the object to be detected in the area to be illuminated according to the preceding continuous multiple frames of images, and according to the hand information and the posture information; and the second target behavior state is determined according to the preceding continuous multiple frames of images;

[0044] The lighting device control module is used to control the lighting state of the lighting device according to the target behavior state.

[0045] In a third aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any lighting control method described in the first aspect are implemented.

[0046] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any lighting control method described in the first aspect are implemented.

[0047] The present application can obtain a plurality of preceding continuous frames of images of an object to be detected; determine the target behavior state of the object to be detected based on the preceding continuous frames of images, and the target behavior state includes a first target behavior state and / or a second target behavior state; wherein, the first target behavior state is determined based on the hand information and posture information of the object to be detected in the area to be illuminated by obtaining the hand information and posture information according to the preceding continuous frames of images; and the second target behavior state is determined based on the preceding continuous frames of images. The target behavior state includes a first target behavior state and / or a second target behavior state, and by determining the target behavior state of the object to be detected, the lighting demand of the object to be detected is accurately judged, and then the lighting state of the lighting device is controlled according to the lighting demand, thereby effectively improving the intelligence of lighting control.

[0048] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0050] Figure 1 is a flow chart of a lighting control method in one embodiment;

[0051] Figure 2 is a flowchart of a lighting control method in a specific embodiment;

[0052] Figure 3 It is a schematic diagram of the lighting control system architecture in a specific embodiment;

[0053] Figure 4 A schematic diagram of an action recognition model architecture in a specific embodiment;

[0054] Figure 5 It is a schematic diagram of the lighting control system architecture in another specific embodiment;

[0055] Figure 6 is a structural block diagram of a lighting control device in one embodiment;

[0056] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the", "these" and the like in this application do not represent quantitative restrictions, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. Usually, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0059] The terms "module", "unit", etc. used below are a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in hardware, the implementation of software, or a combination of software and hardware is also possible and conceivable.

[0060] In the present application embodiment, Figure 1 As shown, a lighting control method is provided, comprising the following steps:

[0061] S201: Acquire a plurality of consecutive frames of images of the object to be detected.

[0062] S203: determining the target behavior state of the object to be detected according to the preceding continuous multiple frames of images, wherein the target behavior state includes a first target behavior state and / or a second target behavior state; wherein the first target behavior state is determined by obtaining hand information and posture information of the object to be detected in the area to be illuminated according to the preceding continuous multiple frames of images, and according to the hand information and the posture information; and the second target behavior state is determined according to the preceding continuous multiple frames of images;

[0063] In the embodiment of the present application, the preceding continuous multiple frames of images can be obtained by continuously shooting images by a camera device, or by shooting a video by a camera device, and then obtaining the preceding continuous multiple frames of images in the video according to a preset number of frames. The preceding continuous multiple frames of images can include any one of an IR image, a depth image, an RGB image, and a grayscale image, or any combination of multiple thereof.

[0064] In order to more accurately determine the lighting requirements of the object to be detected and thus improve the intelligent level of lighting control, the present application determines the target behavior state of the object to be detected based on the preceding continuous multiple frames of images, and then reflects the lighting requirements of the object to be detected through the target behavior state to control the lighting state of the lighting device. The target behavior state of the object to be detected includes a first target behavior state and / or a second target behavior state.

[0065] First, the method for determining the first target behavior state is explained through the embodiments of the present application.

[0066] In the embodiment of the present application, the area to be illuminated may include an area that needs additional brightness, such as a dark area with weak light or an area that needs further additional lighting. In some specific embodiments, the area to be illuminated can be determined by comparing the brightness A1 of the area to be confirmed with the preset brightness A0, and if A1 is not greater than A0, the area to be confirmed is determined to be the area to be illuminated.

[0067] In an embodiment of the present application, the hand information and posture information of the object to be detected can be obtained by obtaining a video or image of the area to be illuminated. In some embodiments, the video or image of the area to be illuminated can be obtained by a camera. If the area to be illuminated is a dark light scene with low brightness, and it is difficult for an RGB lens to capture a valid image, an IR image (Infrared Image) can be obtained by an infrared camera or a depth image can be obtained by a depth camera as the image of the area to be illuminated. Of course, in order to improve the accuracy of obtaining image information, the IR image and the depth image can also be used as the image of the area to be illuminated at the same time, so as to more accurately determine the hand information and posture information of the object to be detected in the image.

[0068] In an embodiment of the present application, the object to be detected may include a target human body in the image of the area to be illuminated. In some embodiments, the hand information includes at least one of the following: position information and motion information of one or more hands, and the posture information includes at least one of the following: line of sight information, head posture information, and human body posture information.

[0069] Inputting the preceding continuous multi-frame image into the detection model can obtain the hand information and posture information of the object to be detected. The detection model may include a hand detection module and a posture detection module, and the detection model may be a deep learning network model. In some specific embodiments, the hand detection module may detect the position information of the hand through a deep learning network, and continuously track the hand through a tracking algorithm, and output the motion information of the hand. In other embodiments, the hand detection module may also output finger motion information. The posture detection module may determine the sight line information through a sight line estimation algorithm; the posture detection module may also determine the human joint point information through a deep learning network, and then calculate the head orientation in the three-dimensional space according to the human joint point to determine the head posture information; the posture detection module may also determine the human posture information through the human joint point information. It should be noted that, since the area to be illuminated may be in a small space, the body of the object to be detected may be blocked by obstacles, so when obtaining the hand information and posture information of the object to be detected, it is not necessarily necessary to obtain the full body image of the object to be detected, and only the necessary hand image information or joint image information is required.

[0070] In the embodiment of the present application, the first target behavior state of the object to be detected can be determined based on the hand information and the posture information.

[0071] In an embodiment of the present application, the first target behavior state includes a state requiring lighting or a state not requiring lighting. The action intention of the object to be detected can be judged based on the hand information and posture information, and then the first target behavior state of the object to be detected is determined. For example, if the area pointed to by at least one of the sight direction, head direction or body direction is the same area as the area where the hand appears, it can be determined that the object to be detected has a lighting demand in this area, and the first target behavior state is determined to be a state requiring lighting. On the other hand, if the area where the hand of the object to be detected appears is different from the area pointed to by the sight direction, head direction and body direction, it means that the action of the object to be detected has no obvious intention, and the first target behavior state of the object to be detected is a state not requiring lighting. The embodiment of the present application can set the judgment conditions for determining the first target behavior state according to different application scenarios.

[0072] In order to further improve the intelligence level of lighting control, the lighting control method provided in the embodiment of the present application also includes: determining one or more sub-areas to be illuminated in the area to be illuminated, wherein the sub-areas to be illuminated include at least one of the following: a preset area for picking up objects and an area for reading objects.

[0073] In the embodiment of the present application, one or more sub-areas to be illuminated can be set in advance in the area to be illuminated, or one or more sub-areas to be illuminated can be divided in real time in the area to be illuminated through real-time image recognition, or the division can be performed by a combination of the two methods.

[0074] In some embodiments, the preset area to be picked up may include an item placement area and a ground area. Different preset areas to be picked up may be set according to different application scenarios. For example, if the application scenario is an in-car scene, the preset area to be picked up may include a preset reading area, a storage box area, a seat back storage bag area, and a door handle area, and the ground area may include the bottom area of ​​the vehicle. In other embodiments, if the application scenario is an indoor scene, the preset area to be picked up may include a locker area, a socket area, etc., and the ground area may include an indoor ground area, etc. In other embodiments, the reading area can be confirmed by identifying the reading material, and the reading material may include books, mobile phones, computers, etc.

[0075] The embodiment of the present application can further improve the intelligence level of lighting control by further dividing the area to be illuminated into one or more sub-areas to be illuminated, and then controlling the lighting states of lighting devices at different positions according to the first target behavior state.

[0076] The following describes a method for determining the second target behavior state through an embodiment of the present application.

[0077] In an embodiment of the present application, the second target behavior state of the object to be detected can be determined by extracting image features of the preceding continuous multi-frame images and performing action recognition on the image features. In some embodiments, the preceding continuous multi-frame images can be input into an action recognition model, and the action recognition model outputs the target recognition action of the object to be detected, and then the second target behavior state is determined according to the target recognition action. The action recognition model is trained based on multi-frame image sample data and preset action recognition training text.

[0078] In other embodiments, the data features extracted by the action recognition model can be used as training data to train at least one action recognition sub-model, and each action recognition sub-model is used to recognize at least one target action of the object to be detected. After the action recognition sub-model is trained, after the previous continuous multiple frame images are input into the action recognition sub-model, the action recognition sub-model outputs the target recognition action of the object to be detected, and then determines the second target behavior state according to the target recognition action.

[0079] It should be noted that the training method of the action recognition model and the action recognition sub-model can refer to the training method of the model in the traditional technology, and this application will not go into details.

[0080] After determining the target recognition action of the object to be detected, the second target behavior state can be determined accordingly. The second target behavior state includes a state requiring lighting or a state not requiring lighting. For example, if the target recognition action is determined to be "reading a book" or "picking up an item", it can be determined that the object to be detected has a lighting demand, and the second target behavior state is determined to be a state requiring lighting. If the target recognition action is determined to be "adjusting sitting posture", etc., it means that the action of the object to be detected has no obvious intention, and the second target behavior state of the object to be detected is a state not requiring lighting. The embodiment of the present application can set the judgment conditions for determining the second target behavior state according to different application scenarios.

[0081] The following describes a method for determining a target behavior state through an embodiment of the present application.

[0082] In some embodiments, the target behavior state includes a first target behavior state, that is, the first target behavior state can be used as the target behavior state, as the behavior state of the object to be detected that is finally determined. In other embodiments, the target behavior state includes a second target behavior state, that is, the second target behavior state can be used as the target behavior state. Specifically, since the first target behavior state is mainly determined based on hand information and posture information, the first target behavior state can be used as the target behavior state in scenarios where lighting needs to be judged based on hand information or posture information, such as reading books, picking up items, etc., to improve the accuracy of the judgment of the lighting needs of the object to be detected. In contrast, for other conventional lighting needs judgment scenarios, the second target behavior state can be used as the target behavior state to judge the lighting needs by identifying preset actions.

[0083] In other embodiments, the target behavior state also includes a first target behavior state and a second target behavior state, which is explained below through the embodiments of the present application.

[0084] In some embodiments, if the determined first target behavior state and the second target behavior state are different, the first target behavior state can be repeatedly determined for a first preset number of times, and the second target behavior state can be repeatedly determined for a second preset number of times according to the method of step S203 in the above embodiment. Among the first target behavior state for the first preset number of times and the second target behavior state for the second preset number of times, if the number of times the lighting-required state is determined is larger, the target behavior state is determined to be the lighting-required state; conversely, if the number of times the lighting-unrequired state is determined is larger, the target recognition behavior state is determined to be the lighting-required state; if the number of times the lighting-required state is determined is the same as the lighting-unrequired state, the lighting-required state and the lighting-unrequired state are randomly selected as the target behavior state.

[0085] In other embodiments, when determining the first target behavior state, the preceding multiple continuous frames of images can be input into the detection model to obtain the hand information and posture information of the object to be detected, and the detection model can also output a first confidence level, and the first confidence level is used as the confidence level that matches the first target behavior state. When determining the second target behavior state, the preceding multiple continuous frames of images can be input into the action recognition model or the action recognition sub-model, and then the target recognition action of the object to be detected is output, and the action recognition model or the action recognition sub-model can also output a second confidence level, and the second confidence level is used as the confidence level that matches the second target behavior state. In some specific embodiments, if the first confidence level is greater than the second confidence level, the first target behavior state is used as the target behavior state; if the first confidence level is less than the second confidence level, the second target behavior state is used as the target behavior state. In other specific embodiments, a first weight can also be set for the first target behavior state, and a second weight can be set for the second target behavior state. If the product of the first weight and the first confidence is greater than the product of the second weight and the second confidence, the first target behavior state is taken as the target behavior state; if the product of the first weight and the first confidence is less than the product of the second weight and the second confidence, the second target behavior state is taken as the target behavior state.

[0086] S207: Controlling the lighting state of the lighting device according to the target behavior state.

[0087] In an embodiment of the present application, the lighting device may include a lighting device covering the entire area to be illuminated and / or a lighting device covering a sub-area to be illuminated. In some embodiments, if the target behavior state is a lighting-required state, the lighting device covering the entire area to be illuminated may be controlled to turn on, and / or the lighting device of the sub-area to be illuminated that matches the target behavior state may be controlled to turn on. In other embodiments, if the target behavior state is a lighting-required state, the lighting device covering the entire area to be illuminated may be controlled to turn off, and / or the lighting device of the sub-area to be illuminated that matches the target behavior state may be controlled to turn off. In some specific embodiments, the lighting device may include a vehicle-mounted reading light or a vehicle-mounted lighting light, and the lighting device may be arranged at a position such as the roof and / or the A-pillar.

[0088] The present application can obtain a plurality of preceding continuous frames of images of an object to be detected; determine the target behavior state of the object to be detected based on the preceding continuous frames of images, and the target behavior state includes a first target behavior state and / or a second target behavior state; wherein, the first target behavior state is determined based on the hand information and posture information of the object to be detected in the area to be illuminated by obtaining the hand information and posture information according to the preceding continuous frames of images; and the second target behavior state is determined based on the preceding continuous frames of images. The target behavior state includes a first target behavior state and / or a second target behavior state, and by determining the target behavior state of the object to be detected, the lighting demand of the object to be detected is accurately judged, and then the lighting state of the lighting device is controlled according to the lighting demand, thereby effectively improving the intelligence of lighting control.

[0089] The following is an explanation of the implementation method of determining the first target behavior state of the object to be detected through the examples of the present application.

[0090] In some embodiments, the first target behavior state is determined according to the following method:

[0091] S401: Determine whether there is a hand in the sub-area to be illuminated according to the hand information.

[0092] S403: When there is no hand in the sub-area to be illuminated, the first target behavior state is that no illumination is required.

[0093] S405: When there is a hand in the sub-area to be illuminated, determining a first target behavior state of the object to be detected according to the posture information.

[0094] In the embodiment of the present application, it is possible to directly or indirectly determine whether a hand is in the sub-area to be illuminated based on the hand information. In some embodiments, the direct determination method may include obtaining the current position of the hand based on the hand position information, and then determining whether the hand is in the sub-area to be illuminated; the indirect determination method may include determining whether the current hand has a tendency to move to the sub-area to be illuminated based on the hand motion information. Even if the hand is blocked in the currently acquired image frame, it is possible to determine whether the current hand is in the sub-area to be illuminated by an indirect determination method. Specifically, the real-time hand motion trajectory information may be obtained based on the hand motion information, and it is determined whether the current hand has moved to the sub-area to be illuminated based on the real-time hand motion trajectory information.

[0095] When there is no hand in the sub-area to be illuminated, it means that the object to be detected has no lighting demand in the sub-area to be illuminated, and the first target behavior state is determined as no lighting required. When there is a hand in the sub-area to be illuminated, the first target behavior state of the object to be detected is determined according to the posture information.

[0096] Since the hand information can clearly reflect the action intention of the object to be detected, by judging whether there is a hand in the sub-area to be illuminated, it can be more accurately judged whether the object to be detected has lighting needs. On the other hand, in order to further improve the intelligence of lighting control, when there is a hand in the sub-area to be illuminated, the first target behavior state of the object to be detected can be further determined in combination with the posture information.

[0097] In some embodiments, determining the first target behavior state of the object to be detected according to the posture information includes:

[0098] S501: Determine a sight line direction according to the sight line information.

[0099] S503: Determine the head direction according to the head posture information; and / or determine the body direction according to the human body posture information.

[0100] S505: Determine the first target behavior state according to the sight direction, the head direction and / or the body direction.

[0101] In the embodiment of the present application, the line of sight direction of the object to be detected can be determined based on the line of sight information; the head direction of the object to be detected can be determined based on the head posture information; the body direction of the object to be detected can be determined based on the human body posture information. The first target behavior state can then be determined based on at least one of the line of sight direction, head direction, and body direction. In actual application scenarios, the judgment conditions for the first target behavior state can be set according to needs.

[0102] In some embodiments, determining the first target behavior state based on the line of sight direction, the head direction and / or the body direction includes: when at least one of the line of sight direction, the head direction and the body direction of the object to be detected points to the sub-area to be illuminated where the hand is located, determining that the first target behavior state is a state requiring illumination.

[0103] In the embodiment of the present application, the sight direction, head direction and / or body direction of the object to be detected points to the sub-area to be illuminated where the hand is located, indicating that the object to be detected may need to pick up items in the sub-area to be illuminated, or indicating that the object to be detected may have the intention to read the reading material in the sub-area to be illuminated, and the first target behavior state can be determined to be the state requiring illumination. For example, in a vehicle, when the sight direction, head direction, and body direction of the rear passenger all point to the sub-area to be illuminated where the hand is located, it indicates that the sub-area to be illuminated may be a preset area to be picked up, and the preset area to be picked up may be located behind the front seat back or at the bottom of the vehicle. At this time, it can be determined that the object to be detected has the intention to pick up the object, and the first target behavior state is determined to be the state requiring illumination.

[0104] In some embodiments, determining the first target behavior state based on the line of sight direction, the head direction and / or the body inclination direction also includes: when the line of sight direction, head direction and body direction of the object to be detected are not pointing to the sub-area to be illuminated where the hand is located, determining that the first target behavior state is a state where no lighting is required.

[0105] In the embodiment of the present application, if only the hand is in the sub-area to be illuminated, and the sight direction, head direction and body direction of the object to be detected are not pointing to the sub-area to be illuminated where the hand is, it means that the hand of the object to be detected may appear in the sub-area to be illuminated meaninglessly, or it means that the object to be detected is familiar with the location of the object to be picked up, or it may be that the object to be detected has completed the action of taking, picking up or reading, so it can be determined that the first target behavior state of the object to be detected is the state without illumination. For example, when the object is successfully picked up from the ground, the body direction and head direction of the object to be detected do not point to the sub-area to be illuminated, and the sight line does not point to the sub-area to be illuminated, which indicates that the behavior of picking up the object is completed and the lighting device can be turned off.

[0106] In an embodiment of the present application, by further judging the action intention of the object to be detected by at least one of the line of sight direction, head direction, and body direction, it is possible to more accurately judge whether the object to be detected has lighting needs, thereby effectively improving the intelligence level of lighting control.

[0107] In some embodiments, when the sub-area to be illuminated where the hand is located is the reading material area, determining the first target behavior state based on the line of sight direction, the head direction and / or the body direction also includes: when the line of sight direction and the head direction of the object to be detected are not pointing to the sub-area to be illuminated where the hand is located, and the duration is not less than a first time threshold, determining that the first target behavior state is a state where no illumination is required.

[0108] Specifically, when the sub-area to be illuminated where the hand is located is a reading area, the following situations may occur:

[0109] The sight direction, head direction and / or body direction of the subject to be detected points to the sub-area to be illuminated where the hand is located, indicating that the subject to be detected may have an intention to read the reading material in the sub-area to be illuminated, and the first target behavior state can be determined as the state requiring illumination;

[0110] The sight direction, head direction and body direction of the subject to be detected are not pointing to the sub-area to be illuminated where the hand is located, indicating that the subject to be detected has completed reading or is not reading. At this time, the first target behavior state is determined to be a state where no illumination is required;

[0111] When the line of sight and head direction of the object to be detected do not point to the sub-area to be illuminated where the hand is located, and the duration is not less than the first time threshold, it means that the object to be detected has completed reading or has not read. At this time, the first target behavior state is determined to be a state without lighting.

[0112] The following is an explanation of the implementation method of determining the second target behavior state of the object to be detected through the examples of the present application.

[0113] In some embodiments, the second target behavior state is determined according to the following method:

[0114] S601: Input the preceding continuous multiple frame images into an action recognition model, and output a target recognition action of the object to be detected, wherein the action recognition model includes a text encoder, a video encoder, and a multi-task decoder, wherein the text encoder is used to extract text features of a preset action recognition text, the video encoder is used to extract image features of the preceding continuous multiple frame images, and the multi-task decoder is used to determine the target recognition action based on the text features and the image features.

[0115] S603: Determine the second target behavior state according to the target recognition action.

[0116] In the embodiment of the present application, the action recognition model includes a large multimodal model, such as CLIP (Contrastive Language-Image Pre-Training, visual language multimodal model), etc. The action recognition model can be used to identify the actions of the object to be detected, such as reading, picking up objects, etc. Compared with the traditional deep learning model, the multi-model large model has stronger expression and generalization capabilities, and can learn more feature information of video images. The multimodal large model achieves cross-modal understanding by learning how to compare text features and image features. The action recognition model in the embodiment of the present application can introduce time series for modeling and fine-tune the multimodal large model.

[0117] In some embodiments, the training set of the action recognition model may include video sequence samples and action text description sample pairs. Specifically, m frames of images can be randomly selected from a continuous N-frame action image sequence as input video sequence training sample data, where m is a preset value, and the training sample data is subjected to data disturbances such as cropping, adding noise, and changing image quality, thereby simulating a variety of real-life scenarios. The above data enhancement operation enables the action recognition model to better adapt to the action recognition tasks of the objects to be detected under different conditions.

[0118] Comparative learning training of the action recognition model requires a large amount of corresponding video image data and text data, and has high requirements for computing hardware. Therefore, fine-tuning can be performed on a small data set for the actual action recognition task based on the pre-trained model. Specifically, when training the action recognition model in the embodiment of the present application, the pre-trained model parameters can be fine-tuned on the data sample set of passenger behavior.

[0119] In an embodiment of the present application, the action recognition model includes a text encoder, a video encoder, and a multi-task decoder. The text encoder is used to extract and encode the preset action recognition text information to obtain text features. The video encoder is used to extract and encode the preceding continuous multi-frame image information to obtain image features. The multi-task decoder is used to convert multimodal features into the same feature space for expression, to achieve alignment of different modal features of text features and image features, and to determine the target recognition action based on the text features and image features. In some instances, the action recognition model is trained using a contrastive learning strategy, and the model is trained using natural language as a supervisory signal, so that the image features and text features generated by it have multimodal characteristics.

[0120] In the embodiment of the present application, the second target behavior state is determined according to the target recognition action. The description of step S203 in the above embodiment can be referred to and will not be repeated here.

[0121] In some embodiments, the determining the target recognition action according to the text feature and the image feature includes:

[0122] S701: Classify the text features of the preset action recognition text to determine the action category label.

[0123] S703: Matching the action category label with the image feature based on contrastive learning, and determining the target recognition action according to the matching result.

[0124] In the embodiment of the present application, the target recognition action can be determined in the form of "zero sample" or few samples to achieve "infinitely fine-grained" label classification. "Zero sample" learning enables the action recognition model to infer and identify new behavior categories through natural language prompts (text descriptions) when the specific category is unknown. This method is based on the characteristics of large multimodal models (such as CLIP), and allows the action recognition model to perform action recognition in the case of unknown data categories through the alignment of text features and image features.

[0125] The text features of the preset action recognition text are classified to determine the action category label, such as reading action, picking action. The action category label is matched with the image feature based on contrastive learning, and the target recognition action is determined according to the matching result. Specifically, in the multimodal feature space, the image feature and the action category label are compared and matched. If the action of the object to be detected is consistent with the action category label, the matching result can be determined as the corresponding action category, and then the target recognition action of the object to be detected is determined.

[0126] In some embodiments, the second target behavior state is further determined according to the following method:

[0127] S801: Input the preceding multiple consecutive frames of images into an action recognition sub-model, and output a target recognition action of the object to be detected, wherein the action recognition sub-model is trained based on the text features and the image features.

[0128] In the embodiment of the present application, the common features extracted by the multi-model large model are used as a training set to train an action recognition sub-model for processing a specific action recognition task. For example, a certain action recognition sub-model can be used to recognize reading actions, another action recognition sub-model can be used to recognize picking actions, and so on.

[0129] In the embodiment of the present application, the target recognition action is determined by the action recognition sub-model, which has higher flexibility and accuracy. Especially in the case of labeled data, the action recognition sub-model can be fully optimized and the action recognition task can be efficiently decomposed. In addition, the computational cost of the action recognition sub-model is low, and the parameters can be fine-tuned for a certain type of action recognition task, thus solving the problem of difficulty in fine-tuning a large multimodal model in a specific action recognition task.

[0130] In some embodiments, controlling the lighting state of the lighting device according to the first target behavior state and / or the second target behavior state includes:

[0131] S901: Controlling the lighting state of the sub-lighting device corresponding to the sub-area to be illuminated according to the target behavior state.

[0132] S903: When the target behavior state is a lighting-required state, control the sub-lighting device corresponding to the sub-area to be illuminated to be in an on state.

[0133] S905: When the target behavior state is a state where lighting is not required, controlling the sub-lighting device corresponding to the sub-area to be illuminated to be in a closed state.

[0134] By controlling the lighting state of the sub-lighting device corresponding to the sub-area to be illuminated according to the target behavior state, lighting can be provided for the sub-area to be illuminated in a targeted manner, thereby realizing intelligent zoned light control. When the target behavior state is a lighting-required state, the sub-lighting device corresponding to the sub-area to be illuminated is controlled to be in an on state; when the target behavior state is a lighting-unrequired state, the sub-lighting device corresponding to the sub-area to be illuminated is controlled to be in an off state. The present application can effectively improve the accuracy of the judgment of the lighting demand of the object to be detected, thereby improving the intelligence level of lighting control, by flexibly setting the judgment method of the target behavior state and controlling the lighting state of the lighting device of the lighting sub-area that matches the target behavior state.

[0135] The following uses the in-car application scenario as a specific embodiment to illustrate the lighting control method provided by the present application. In this implementation, the first target behavior state is used as the target behavior state, such as Figure 2 As shown, the brightness of the area to be illuminated is compared with the preset brightness. If the brightness of the area to be illuminated is not greater than the preset brightness, after obtaining the previous continuous frames, one or more sub-areas to be illuminated are determined in the area to be illuminated, and passenger information, such as hand information and posture information, is detected. If there is no hand in the area to be illuminated, the lighting device corresponding to the sub-area to be illuminated is controlled to be in a closed state. If there is a hand in the area to be illuminated, that is, the hand position is in the sub-area to be illuminated or the hand movement trajectory predicts that the current hand is in the sub-area to be illuminated, then combined with other information such as body posture information, it is determined whether any of the sight, head, and body points to the sub-area to be illuminated corresponding to the hand. If so, the lighting device corresponding to the sub-area to be illuminated is in an open state, that is, the lighting device is turned on or the lighting device is kept in a continuously turned-on state. Further, if the sight, head, and body in the body posture information do not point to the sub-area to be illuminated corresponding to the hand, and / or the sub-area to be illuminated is a reading sub-area and the sight or head leaves the sub-area to be illuminated for more than a certain period of time, the lighting device corresponding to the sub-area to be illuminated is controlled to be in a closed state, that is, the lighting device is turned off or the lighting device is kept in a continuously turned-off state.

[0136] The lighting control system provided by the present application is described below through a specific embodiment of the present application. In this implementation, the first target behavior state is used as the target behavior state. The lighting control system architecture is as follows Figure 3 As shown, after obtaining the IR video stream or the depth video stream, it is input into the detection module. After detecting and tracking any hand shape and multi-hand movements, estimating the gaze and head posture, and detecting the reading material, the hand information and posture information of the object to be detected are determined by the analysis and judgment module. Finally, the target behavior state of the object to be detected is determined by locating the lighting area and the judgment result of the analysis and judgment module. If the action of the target to be detected is to pick up an object, the lighting area is controlled to turn on for a short time; if the action of the target to be detected is to read, the lighting area is controlled to turn on for a long time.

[0137] The following describes the action recognition method of the action recognition model through the specific embodiments of the present application. Figure 4 The action recognition model architecture is shown in the figure. The action recognition model includes a text encoder, a video encoder, and a multi-task decoder. The text encoder is used to extract text features of a preset action recognition text. The preset action recognition text includes text descriptions such as "reading a book", "looking at a pad", "picking up an object on the ground" or other behaviors. The video encoder is used to extract image features of the preceding continuous multiple frames of images in the input video sequence. The multi-task decoder is used to output the target recognition action based on the text features and the image features.

[0138] The lighting control system provided by the present application is described below through another specific embodiment of the present application. In this implementation, the target behavior state is determined by the first target behavior state and the second target behavior state. Figure 5 As shown, after obtaining the video stream, the detection module and the multimodal large model are input respectively. After the detection module detects and tracks any hand shape and multi-hand actions, estimates the posture of the line of sight and the head, and detects the reading material, the hand information and posture information of the object to be detected are determined by the analysis and judgment module. The text prompt words used to describe the action are also input into the multimodal large model, and the text features of the text prompt words are extracted by the multimodal large model, and the image features of the video stream are extracted. After matching the text features with the image features, the target recognition action is output. Finally, by locating the lighting area, analyzing the judgment results of the judgment module and the target recognition action, the target behavior state of the object to be detected is jointly determined. If the action of the target to be detected is to pick up an object, the light of the area to be illuminated is controlled to be turned on for a short time; if the action of the target to be detected is to read, the light of the area to be illuminated is controlled to be turned on for a long time.

[0139] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0140] Based on the same inventive concept, the embodiment of the present application also provides a lighting control device 1200 for implementing the lighting control method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more lighting control device embodiments provided below can refer to the limitations of the lighting control method above, and will not be repeated here.

[0141] In one embodiment, Figure 6 As shown, a lighting control device 1200 is provided, comprising:

[0142] The image acquisition module 1201 is used to acquire a plurality of consecutive frames of images of the object to be detected;

[0143] The target behavior state determination module 1202 is used to determine the target behavior state of the object to be detected according to the preceding continuous multiple frames of images, wherein the target behavior state includes a first target behavior state and / or a second target behavior state; wherein the first target behavior state is determined by obtaining hand information and posture information of the object to be detected in the area to be illuminated according to the preceding continuous multiple frames of images, and according to the hand information and the posture information; and the second target behavior state is determined according to the preceding continuous multiple frames of images;

[0144] The lighting device control module 1204 is used to control the lighting state of the lighting device according to the target behavior state.

[0145] In one embodiment, the device 1200 further includes a device for determining a sub-area to be illuminated, which is used to determine one or more sub-areas to be illuminated within the area to be illuminated, wherein the sub-areas to be illuminated include at least one of the following: a preset area for objects to be picked up and an area for reading objects.

[0146] In one embodiment, the hand information includes at least one of the following: position information and motion information of one or more hands, and the posture information includes at least one of the following: line of sight information, head posture information, and human body posture information.

[0147] In one embodiment, the target behavior state determination module 1202 is also used to determine whether there is a hand in the sub-area to be illuminated based on the hand information; when there is no hand in the sub-area to be illuminated, the first target behavior state is a state where no illumination is required; when there is a hand in the sub-area to be illuminated, the first target behavior state of the object to be detected is determined based on the posture information.

[0148] In one embodiment, the target behavior state determination module 1202 is also used to determine the line of sight direction based on the line of sight information; determine the head direction based on the head posture information; and / or determine the body direction based on the human body posture information; and determine the first target behavior state based on the line of sight direction, the head direction and / or the body direction.

[0149] In one embodiment, the target behavior state determination module 1202 is also used to determine that the first target behavior state is a state requiring lighting when at least one of the sight direction, head direction and body direction of the object to be detected points to the sub-area to be illuminated where the hand is located.

[0150] In one embodiment, the target behavior state determination module 1202 is also used to determine that the first target behavior state is a state that does not require lighting when the sight direction, head direction and body direction of the object to be detected do not point to the sub-area to be illuminated where the hand is located.

[0151] In one embodiment, when the sub-area to be illuminated where the hand is located is the reading material area, the target behavior state determination module 1202 is also used to determine that the first target behavior state is a state where no illumination is required when the line of sight and the head direction of the object to be detected are not pointing to the sub-area to be illuminated where the hand is located, and the duration is not less than the first time threshold.

[0152] In one embodiment, the target behavior state determination module 1202 is also used to input the preceding continuous multiple frame images into an action recognition model, and output the target recognition action of the object to be detected. The action recognition model includes a text encoder, a video encoder and a multi-task decoder. The text encoder is used to extract text features of a preset action recognition text, the video encoder is used to extract image features of the preceding continuous multiple frame images, and the multi-task decoder is used to determine the target recognition action based on the text features and the image features; and determine the second target behavior state based on the target recognition action.

[0153] In one embodiment, the target behavior state determination module 1202 is also used to classify the text features of the preset action recognition text to determine the action category label; match the action category label with the image feature based on contrast learning, and determine the target recognition action according to the matching result.

[0154] In one embodiment, the target behavior state determination module 1202 is also used to input the previous continuous multiple frame images into an action recognition sub-model, and output the target recognition action of the object to be detected. The action recognition sub-model is trained based on the text features and the image features.

[0155] In one embodiment, the lighting device control module 1204 is also used to control the lighting state of the sub-lighting device corresponding to the sub-area to be illuminated according to the target behavior state; when the target behavior state is a lighting required state, the sub-lighting device corresponding to the sub-area to be illuminated is controlled to be in an on state; when the target behavior state is a lighting not required state, the sub-lighting device corresponding to the sub-area to be illuminated is controlled to be in an off state.

[0156] Each module in the lighting control device 1200 can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0157] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a lighting control method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0158] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0159] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the lighting control method described in any of the above embodiments are implemented.

[0160] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the lighting control method described in any of the above embodiments are implemented.

[0161] In one embodiment, a computer program product is provided, including a computer program, which implements the steps of the lighting control method described in any of the above embodiments when executed by a processor.

[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0163] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0164] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0165] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A lighting control method, characterized in that: The method comprises: Acquire a plurality of consecutive frames of images of the object to be detected; Determining a target behavior state of the object to be detected according to the preceding continuous multiple frames of images, wherein the target behavior state includes a first target behavior state and / or a second target behavior state; Wherein, the first target behavior state is determined by obtaining the hand information and posture information of the object to be detected in the area to be illuminated according to the preceding continuous multiple frames of images, and according to the hand information and the posture information; The second target behavior state is determined according to the preceding continuous multiple frames of images; According to the target behavior state, the lighting state of the lighting device is controlled.

2. The method according to claim 1, characterized in that The method further comprises: One or more sub-areas to be illuminated are determined in the area to be illuminated, wherein the sub-areas to be illuminated include at least one of the following: a preset area for objects to be picked up and an area for reading objects.

3. The method according to claim 2, characterized in that The hand information includes at least one of the following: position information and movement information of one or more hands, and the posture information includes at least one of the following: line of sight information, head posture information, and human body posture information.

4. The method according to claim 3, characterized in that The first target behavior state is determined according to the following method: According to the hand information, determining whether there is a hand in the sub-area to be illuminated; When there is no hand in the sub-area to be illuminated, the first target behavior state is a state where no illumination is required; When there is a hand in the sub-area to be illuminated, a first target behavior state of the object to be detected is determined according to the posture information.

5. The method according to claim 4, characterized in that The determining the first target behavior state of the object to be detected according to the posture information includes: Determining a sight direction according to the sight information; Determining the head direction according to the head posture information; and / or Determining a body direction according to the human body posture information; The first target behavior state is determined according to the sight direction, the head direction and / or the body direction.

6. The method according to claim 5, characterized in that The determining the first target behavior state according to the sight direction, the head direction and / or the body direction includes: When at least one of the sight line direction, the head direction and the body direction of the object to be detected points to the sub-region to be illuminated where the hand is located, it is determined that the first target behavior state is a state requiring illumination.

7. The method according to claim 6, characterized in that The determining the first target behavior state according to the sight direction, the head direction and / or the body inclination direction also includes: When the sight direction, the head direction and the body direction of the object to be detected do not point to the sub-area to be illuminated where the hand is located, it is determined that the first target behavior state is a state where no illumination is required.

8. The method according to claim 6, characterized in that When the sub-area to be illuminated where the hand is located is the reading area, the determining the first target behavior state according to the sight direction, the head direction and / or the body direction further includes: When the sight line direction and the head direction of the object to be detected do not point to the sub-area to be illuminated where the hand is located, and the duration is not less than the first time threshold, it is determined that the first target behavior state is a state that does not require illumination.

9. The method according to claim 1, characterized in that: The second target behavior state is determined according to the following method: Input the preceding continuous multi-frame images into an action recognition model, and output a target recognition action of the object to be detected, wherein the action recognition model includes a text encoder, a video encoder, and a multi-task decoder, wherein the text encoder is used to extract text features of a preset action recognition text, the video encoder is used to extract image features of the preceding continuous multi-frame images, and the multi-task decoder is used to determine the target recognition action according to the text features and the image features; The second target behavior state is determined according to the target recognition action.

10. The method according to claim 9, characterized in that The determining the target recognition action according to the text feature and the image feature comprises: Classifying the text features of the preset action recognition text to determine the action category label; The action category label is matched with the image feature based on contrastive learning, and the target recognition action is determined according to the matching result.

11. The method according to claim 9, characterized in that The second target behavior state is also determined according to the following method: The preceding multiple consecutive frames of images are input into an action recognition sub-model, and a target recognition action of the object to be detected is outputted. The action recognition sub-model is trained based on the text features and the image features.

12. The method according to any one of claims 2 to 11, characterized in that: The controlling the lighting state of the lighting device according to the target behavior state includes: According to the target behavior state, controlling the lighting state of the sub-lighting device corresponding to the sub-area to be illuminated; When the target behavior state is a lighting-required state, controlling the sub-lighting device corresponding to the sub-area to be illuminated to be in an on state; When the target behavior state is a state where lighting is not required, the sub-lighting device corresponding to the sub-area to be illuminated is controlled to be in a closed state.

13. A lighting control device, characterized in that: The device comprises: An image acquisition module is used to acquire a plurality of consecutive frames of images of the object to be detected; A target behavior state determination module, used to determine the target behavior state of the object to be detected according to the preceding continuous multiple frames of images, wherein the target behavior state includes a first target behavior state and / or a second target behavior state; wherein the first target behavior state is determined by obtaining hand information and posture information of the object to be detected in the area to be illuminated according to the preceding continuous multiple frames of images, and according to the hand information and the posture information; and the second target behavior state is determined according to the preceding continuous multiple frames of images; The lighting device control module is used to control the lighting state of the lighting device according to the target behavior state.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.