Method for updating multi-task joint model, control method, device and medium
Patent Information
- Application Number
- CN202610796233.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-25
AI Technical Summary
这类跨任务输出头的预测冲突,在单一任务的评测中往往难以被发现,进而影响多任务联合模型的感知精度与决策可靠性,最终可能导致可移动设备出现运行安全隐患
[0016]基于本公开实施例提供的多任务联合模型的更新方法,通过引入多任务联合模型的多个任务输出头中相关联的两个任务输出头在执行各自目标任务时的兼容性约束,自动检测该两个任务输出头所输出的目标任务结果之间的输出冲突并量化为第一自洽损失值,将第一自洽损失值与多任务联合模型的基本损失值相结合,共同更新多任务联合模型的模型参数,实现了冲突检测与模型训练的闭环协同,无需依赖人工标注数据或经验调参,显著提升了模型迭代效率、降低了成本。
Smart Images

Figure CN122817862A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of model update technology, and more specifically, to an update method, control method, device, and medium for a multi-task joint model. Background Technology
[0002] Multi-task joint models typically include multiple task output heads to simultaneously output prediction results for different tasks, such as vehicle detection, curb detection, parking space recognition, and passable area detection. These models are widely used in perception and decision-making tasks of mobile devices (such as autonomous vehicles), providing crucial data support for the autonomous operation of these devices.
[0003] However, the prediction results output by different task output heads may conflict with each other. For example, the vehicle bounding box output by the vehicle detection head and the curb point sequence output by the curb detection head may have a spatial conflict of "vehicle colliding with the curb"; the red light signal output by the traffic light detection head and the vehicle behavior output by the planning and control head may have a logical conflict of "vehicles proceeding when the light is red". Such prediction conflicts across task output heads are often difficult to detect in the evaluation of a single task, thus affecting the perception accuracy and decision reliability of the multi-task joint model, and may ultimately lead to operational safety hazards for mobile devices.
[0004] In existing technologies, resolving the aforementioned conflicts mainly relies on manually labeled joint training data or algorithm engineers tuning parameters based on experience. However, this approach suffers from problems such as the separation of evaluation and training, long iteration cycles, and high costs. Therefore, how to efficiently detect output conflicts between different output heads in a multi-task joint model and transform the detection results into training optimization signals is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] In view of this, the present disclosure proposes a new technical solution for updating the multi-task joint model.
[0006] According to a first aspect of the present disclosure, a method for updating a multi-task collaborative model is provided, the method comprising: Obtain the target task result set output by the multi-task joint model for the sample data; wherein, the sample data is data generated based on scene data collected by a mobile device, the multi-task joint model includes multiple task output heads, different task output heads correspond to different target tasks, and the target task result set includes the target task results obtained by each of the task output heads performing the corresponding target task; Based on the compatibility constraints of two associated task output heads among the plurality of task output heads when executing their respective target tasks, a first self-consistency loss value is determined between the target task results output by the two task output heads. The model parameters of the multi-task joint model are updated based on the first self-consistent loss value and the basic loss value of the multi-task joint model; wherein the basic loss value is obtained based on the deviation between the target task result set and the true result set corresponding to the sample data.
[0007] Optionally, determining the first self-consistency loss value between the target task results output by the two associated task output heads based on the compatibility constraints of the two associated task output heads when executing their respective target tasks includes: Obtain the first task result output by the first task output head and the second task result output by the second task output head, which are associated among the multiple task output heads; Based on the compatibility constraints of the first task output head and the second task output head when executing their respective target tasks, the first task result and the second task result, a self-consistent evaluation index is generated; wherein, the self-consistent evaluation index is used to determine the degree of conflict between the first task result and the second task result; The first self-consistency loss value between the first task result output by the first task output head and the second task result output by the second task output head is determined based on the self-consistency evaluation index.
[0008] Optionally, determining the first self-consistency loss value between the target task results output by the two associated task output heads based on the compatibility constraints of the two associated task output heads when executing their respective target tasks includes: Obtain the first task result output by the first task output head and the second task result output by the second task output head associated with the plurality of task output heads; Based on the compatibility constraints of the first task output head and the second task output head when executing their respective target tasks, the first task result and the second task result, a self-consistent evaluation index is generated; wherein, the self-consistent evaluation index is used to determine the degree of conflict between the first task result and the second task result; The first self-consistency loss value between the first task result output by the first task output head and the second task result output by the second task output head is determined based on the self-consistency evaluation index.
[0009] Optionally, updating the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model includes: The total loss value of the multi-task joint model is calculated based on the first self-consistent loss value and its corresponding weight coefficient, as well as the basic loss value of the multi-task joint model. The model parameters of the multi-task joint model are updated based on the total loss value.
[0010] Optionally, updating the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model includes: The total loss value of the multi-task joint model is calculated based on the first self-consistent loss value and its corresponding weight coefficient, as well as the basic loss value of the multi-task joint model. The model parameters of the multi-task joint model are updated based on the total loss value.
[0011] Optionally, the method further includes: Obtain the updated multi-task joint model; Obtain the updated set of verification results output by the multi-task joint model for the verification data; The feature value of the compatibility feature is based on the verification results output by the two associated task output headers in the verification result set; If the feature value exceeds the feature evaluation threshold, adjust the weight coefficient of the corresponding first self-consistent loss value, and re-execute the step of calculating the total loss value of the multi-task joint model based on the first self-consistent loss value, its corresponding weight coefficient, and the basic loss value of the multi-task joint model.
[0012] Optionally, the method further includes: Based on the single-task self-consistency constraint of each of the plurality of task output heads when executing the corresponding target task, determine the second self-consistency loss value of the target task result output by each of the task output heads; The step of updating the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model includes: The model parameters of the multi-task joint model are updated based on the second self-consistent loss value, the first self-consistent loss value, and the basic loss value of the multi-task joint model.
[0013] According to a second aspect of the present disclosure, a control method is provided, the method comprising: Acquire real-time input data collected by mobile devices; The real-time input data is input into the multi-task joint model to obtain the real-time task results output by each task output head in the multi-task joint model; wherein, the multi-task joint model is trained according to the training method of the multi-task joint model described in the first aspect above; Motion control is performed on the mobile device based on the real-time task results output by each task output head in the multi-task joint model.
[0014] According to a third aspect of the present disclosure, an electronic device is provided, including a memory and a processor, the memory being configured to store computer instructions, and the processor being configured to invoke the computer instructions from the memory to perform the method as described in the first or second aspect.
[0015] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method as described in the first or second aspect.
[0016] The update method for the multi-task joint model provided in this disclosure introduces compatibility constraints between two related task output heads in the multi-task joint model when executing their respective target tasks. It automatically detects output conflicts between the target task results output by the two task output heads and quantifies them into a first self-consistent loss value. The first self-consistent loss value is combined with the basic loss value of the multi-task joint model to jointly update the model parameters of the multi-task joint model. This achieves closed-loop collaboration between conflict detection and model training, without relying on manually labeled data or experience-based parameter tuning, significantly improving model iteration efficiency and reducing costs.
[0017] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of the present disclosure and, together with their description, serve to explain the principles of the present disclosure.
[0019] Figure 1 This is a schematic diagram of an intelligent connected system to which the methods provided in the embodiments of this disclosure can be applied.
[0020] Figure 2 It is based on Figure 1 The illustrated embodiment provides a schematic diagram of a mobile device.
[0021] Figure 3 This is a flowchart illustrating an update method for a multi-task collaborative model provided in an embodiment of this disclosure.
[0022] Figure 4 This is a flowchart illustrating another method for updating a multi-task collaborative model provided in this embodiment.
[0023] Figure 5This is a flowchart illustrating a control method provided in an embodiment of this disclosure.
[0024] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0025] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0026] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0027] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0028] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0029] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0030] The elements involved in the embodiments of this disclosure may represent part or all of an element. For example, the elements involved in the embodiments of this disclosure may be at least a part of an element or all of an element.
[0031] The elements involved in the embodiments of this disclosure may be one or more, such as "a", "the", "the above", "the", "the foregoing", etc., which are used to indicate that the corresponding element is mentioned for the first time or is mentioned again, and do not have the meaning of limiting the number.
[0032] It should be noted that all actions involving the collection, storage, use, processing, transmission, provision, disclosure, and deletion of data in this disclosure are carried out in accordance with the relevant data protection laws and regulations of the country or region where the data is located, and with the full authorization of the relevant data owner.
[0033] First, the application scenarios of the embodiments of this disclosure will be described.
[0034] Figure 1 This is a schematic diagram of an intelligent connected system 100 to which the methods provided in the embodiments of this disclosure can be applied. Figure 1As shown, the intelligent connected system 100 may include: a mobile device 101, a server 102, and a user terminal 103.
[0035] In some examples, the mobile device 101 can be a mobile device such as a vehicle, robot, ship, or aircraft, for example, a vehicle, ship, or aircraft with a driving automation feature, or an autonomous robot (such as a cargo robot, a probe robot, or a sweeping robot).
[0036] The driving automation function can include advanced driver assistance functions (ADAS) and automated driving functions (AWD). Automated driving, also known as intelligent driving or driverless driving, refers to vehicles equipped with driving automation functions that can perform some or all of the driving tasks, such as environmental perception, decision-making, planning, and control execution. The levels of driving automation functions can refer to the vehicle intelligence classification standards established by the Society of Automotive Engineers (SAE), for example, divided into six levels from L0 to L5. L0 is emergency assistance, L1 is partial driver assistance, L2 is combined driver assistance, L3 is conditional automated driving, L4 is highly automated driving, and L5 is fully automated driving. The above classification of driving automation function levels is merely an example, and this disclosure does not limit the classification standards and levels of driving automation functions.
[0037] In some examples, server 102 can be a single server or a distributed server cluster consisting of multiple servers, and its deployment method can include local servers or cloud servers. Server 102 can communicate with mobile device 101 and / or user terminal 103 via a communication network, providing various services to mobile device 101 and / or user terminal 103. For example, the server can receive sensing data sent by mobile device and provide services such as navigation maps, data analysis, and decision planning to mobile device. Alternatively, the server can receive query commands or control commands sent by user terminal and provide corresponding services to the user.
[0038] In some examples, user terminal 103 can be any form of electronic device providing services to the user, such as a personal computer, laptop, smart tablet, smartphone, smart wearable device, etc. The user can interact with the mobile device or server through the human-computer interaction terminal configured on the mobile device 101, or through user terminal 103. For example, the user can query the status and / or parameters of the mobile device, or control the mobile device to perform set tasks and / or modify configuration parameters, etc. The user terminal runs an application based on the intelligent network system to achieve interaction with the mobile device or server. This application can be a local application, a web application, or a mini-program, etc., and is not limited thereto.
[0039] In some examples, the aforementioned application running on the user's terminal can provide authentication or authorization services to the user. The user who is successfully authenticated and granted the corresponding permissions can query and / or control the mobile device within the scope of the granted permissions.
[0040] The mobile device 101, server 102, and user terminal 103 can communicate via a communication link provided by communication network 104. This communication network 104 can include one or more networks of any type, such as the Internet, Local Area Network (LAN), Wide Area Network (WAN), Virtual Private Network (VPN), Public Switched Telephone Network (PSTN), satellite communication network, Wi-Fi, 2G, 3G, 4G, 5G, 6G, NB-IoT, eMTC, infrared, Bluetooth, NFC, or a combination of these networks. The communication networks between the mobile device 101 and server 102, between the user terminal 103 and server 102, and between the user terminal 103 and mobile device 101 can be the same or different.
[0041] It should be noted that, Figure 1 The structure of the intelligent connected system 100 shown is merely illustrative. The intelligent connected system in this embodiment is not limited to the above structure and may include more or fewer devices as needed, and the devices may be combined or split. For example, the intelligent connected system may not include user terminals and / or servers; as another example, user terminals and servers may be deployed together.
[0042] Figure 2 It is based on Figure 1 The illustrated embodiment provides a schematic diagram of a mobile device 101. As shown... Figure 2As shown, the mobile device 101 may include a sensing component 1011, a computing platform 1012, an execution component 1013, etc. The sensing component 1011, the computing platform 1012, and the execution component 1013 may be connected via a bus or other means.
[0043] In some examples, the sensing component 1011 can be used to collect information about the mobile device itself or externally. The sensing component 1011 may include at least one of a visual sensing unit, radar, positioning and navigation unit, inertial measurement unit (IMU) or other sensing unit. The visual sensor unit may include one or more cameras, the radar may include at least one of lidar, millimeter-wave radar, ultrasonic radar or other radar, and the positioning and navigation unit may include at least one of a GPS system, BeiDou system or other global positioning system.
[0044] In some examples, the computing platform 1012 may include a computing-capable device for processing the sensing information collected by the sensing component 1011 to obtain control information, and sending corresponding control commands to the execution component 1013 to cause the execution component 1013 to perform corresponding actions, thereby realizing the control of the mobile device 101. For example, the computing platform 1012 can perform one or more of the following actions on the mobile device: information collection and processing, positioning, decision-making, planning, and control, thereby realizing the autonomous control of the mobile device. The computing platform 1012 may include at least one processor and at least one memory, wherein each processor can individually or jointly execute instructions stored in the memory to implement the methods provided in the embodiments of this disclosure. The processor in this disclosure embodiment may include at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), Tensor Processing Unit (TPU), Data Processing Unit (DPU), Digital Signal Processor (DSP), Field Programmable Gate Array (FPGA), Programmable Logic Array (PLA), System on Chip (SOC), Application Specific Integrated Circuit (ASIC), Micro Controller Unit (MCU), or other processors. The memory may be implemented using any type of volatile or non-volatile computer-readable storage medium or a combination thereof. In addition to storing instructions, the memory may also store data, such as map data, image data, sound data, text data, configuration parameters of the mobile device, location, orientation, speed, etc. The data stored in the memory can be accessed and used by the processor.
[0045] In some examples, the computing platform of a mobile device can perform computing tasks independently or communicate with a server to complete computing tasks. For example, the computing platform of a mobile device can cooperate with a server to complete corresponding computing tasks. These computing tasks can include any task performed to achieve autonomous control of the mobile device, such as information collection and processing, positioning, decision-making, planning, or control of the mobile device.
[0046] The computing platform 1012 can be located in the mobile device 101. Some or all of the computing platform 1012 can also be located in the server corresponding to the mobile device. For example, some functions of the computing platform 1012 with high real-time requirements can be located in the mobile device, while other functions with low real-time requirements can be located in the server corresponding to the mobile device.
[0047] In some examples, the execution component 1013 is used to perform corresponding actions (such as steering, acceleration, deceleration, braking, etc.) based on the control of the computing platform 1012, so that the mobile device 101 completes the movement task. The execution component 1013 may include, for example, a power component, a braking component, a transmission component, a steering component, etc.
[0048] It should be noted that, Figure 2 The structure of the mobile device 101 shown is merely illustrative. The mobile device in this embodiment is not limited to the above structure and may include more or fewer components as needed. The device may also be combined or disassembled. For example, the mobile device may not include the aforementioned computing platform. Furthermore, the mobile device may also include communication components, interface components, multimedia components, input components, output components, display components, etc.
[0049] In related technologies, multi-task joint models typically include multiple task output heads. The output results between two related task output heads in a multi-task joint model may conflict with each other. Such prediction conflicts across task output heads are often difficult to detect in single-task evaluations. In existing technologies, resolving these conflicts mainly relies on manually annotating joint training data or algorithm engineers tuning parameters based on experience, but this leads to a disconnect between evaluation and training, long iteration cycles, and high costs.
[0050] To address the problems in related technologies, this disclosure provides an update method for a multi-task joint model. Figure 3 This is a flowchart illustrating an update method for a multi-task collaborative model provided in this disclosure. The update method for this multi-task collaborative model can be... Figure 1 The illustrated mobile device and / or server execute. For example... Figure 3 As shown, the update method of the multi-task joint model in this embodiment may include the following steps S310 to S330.
[0051] Step S310: Obtain the target task result set output by the multi-task joint model for the sample data.
[0052] The sample data can be generated from scene data collected by mobile devices. For example, this scene data can be ambient environmental data collected by mobile devices, such as vehicles, using sensors (e.g., cameras, LiDAR, ultrasonic radar, etc.) while in motion; such data includes image data, point cloud data, or ultrasonic data. Based on this scene data, sample data can be generated to update the multi-task joint model.
[0053] The sample data may also include annotation information extracted from the scene data. This annotation data can be ground truth data obtained through manual or automatic annotation, and is used to calculate the basic loss value output by the multi-task joint model.
[0054] The multi-task joint model can be an iterative model to be evaluated. For example, during the update process of the multi-task joint model, the algorithm provider can output an iterative model to be evaluated, and the evaluator can obtain and evaluate this iterative model. This iterative model to be evaluated can be an intermediate version model obtained after several rounds of training, or it can be a candidate model for release.
[0055] This multi-task joint model can include multiple task output heads, and different task output heads can correspond to different target tasks.
[0056] For example, in an autonomous driving scenario, this multi-task joint model can be a bird's-eye view BEV joint training model, including two interconnected task output heads: a vehicle detection head and a curb detection head. The vehicle detection head detects vehicles in the sample data, determines their position and size in the scene, and outputs vehicle bounding box parameters. These parameters can include, for example, the center coordinates, width, height, and rotation angle of the vehicle bounding box, uniquely representing the vehicle's position and size in the scene. The curb detection head performs contour detection and localization of road boundaries (i.e., curbs) in the sample data, determines the spatial distribution of curbs in the scene, and outputs a continuous sequence of curb points to represent the spatial contour morphology of the curbs.
[0057] The vehicle detection head and the curb detection head share the same backbone network of the multi-task joint model, which can execute vehicle detection and curb detection tasks in parallel based on the same set of sample data, and output their respective target task results synchronously, ensuring the spatial correlation of the output results of the two tasks.
[0058] For example, in an automated parking scenario, this multi-task joint model can be a parking joint training model, including two interconnected task output heads: a parking space recognition head and a passable area detection head. The parking space recognition head identifies and locates parking space areas in the sample data, determines the location and range of available parking spaces, and outputs parking space polygon parameters. These polygon parameters can include, for example, the coordinates of the four corner points of the parking space. The parking space polygon parameters uniquely determine the location and outline shape of the parking space in the scene. The passable area detection head detects and segments passable areas in the sample data, determines the area where vehicles can safely drive, and outputs a passable area semantic mask. This semantic mask can be, for example, a binary mask matching the size of the input scene data. Pixel values indicate whether the corresponding location belongs to a passable or impassable area. Impassable areas can include obstacles, curbs, occupied parking spaces, etc., thereby determining the space in the scene where vehicles are prohibited from entering.
[0059] The parking space recognition head and the passable area detection head share the same backbone network of the multi-task joint model, which can execute parking space recognition and passable area detection tasks in parallel based on the same set of sample data, and output their respective target task results synchronously, ensuring spatial consistency and logical compatibility between the parking space area and the passable area.
[0060] The target task result set can include the target task results obtained by each task output head performing the corresponding target task. The multi-task joint model possesses multi-task parallel inference capabilities. After sample data is input into the multi-task joint model, the backbone network extracts unified scene features and distributes them to each task output head. Each task output head independently completes task inference based on shared features and outputs target task results adapted to its own task type.
[0061] For example, continuing with the multi-task joint model as the bird's-eye view BEV joint training model, after inputting sample data into the multi-task joint model, the vehicle detection head can output vehicle bounding box parameters, and the roadside detection head can output a continuous sequence of roadside points. The above outputs together constitute the target task result set.
[0062] For example, continuing with the multi-task joint model as the parking joint training model, after the sample input data is input into the multi-task joint model, the parking space recognition head can output the parking space polygon parameters, and the passable area detection head can output the binary mask. The above outputs together constitute the target task result set.
[0063] It should be noted that there may be conflicts or incompatibilities between different target task results in the target task result set. For example, the vehicle bounding box output by the vehicle detection head and the curb point sequence output by the curb detection head may have a geometric conflict of "vehicle hitting the curb"; the passable area output by the passable area detection head and the lane line output by the lane line detection head may have a logical conflict of "passable area and lane line are inconsistent". Such cross-output head conflict problems are the main technical problems that this disclosure embodiment aims to solve.
[0064] Step S320: Based on the compatibility constraints of two related task output heads among the multiple task output heads when executing their respective target tasks, determine the first self-consistent loss value between the target task results output by the two task output heads.
[0065] In some examples, compatibility constraints may include at least one of the following: physical space constraints, causal relationship constraints, and application scenario business constraints.
[0066] Among them, physical space constraints can be used to constrain the spatial position matching relationship between the target task results output by different task output heads.
[0067] Continuing with the example of the multi-task joint model as the bird's-eye view BEV joint training model, the vehicle bounding box output by the vehicle detection head and the roadside point sequence output by the roadside detection head must meet the physical space constraints: the area where the vehicle is located must not spatially intrude into the roadside contour.
[0068] Continuing with the multi-task joint model as the parking joint training model, the parking space polygon area output by the parking space recognition head and the binary mask output by the passable area detection head must meet the physical space constraints: the parking space area should be located within the passable area, and the parking space should not overlap with the curb area.
[0069] Among these constraints, causal relationship constraints can be used to constrain the causal logical consistency between the target task results output by different task output heads. For example, if the obstacle detection head outputs the motion state as "stationary," then the predicted displacement output by the trajectory prediction head should not be too large (i.e., stationary objects should not produce significant displacement); if the obstacle detection head outputs the motion state as "high-speed movement," then the predicted displacement output by the trajectory prediction head should not be too small (i.e., high-speed moving objects should not have too small displacement). Through this constraint, output conflicts that violate causal logic can be detected and penalized.
[0070] The application scenario business constraints can be used to constrain the matching relationship between the target task results output by different task output heads and preset business rules. For example, if the traffic light detection head outputs a "red light" status, the vehicle passage behavior output by the planning control head should be "stop," not "pass"; if the traffic light detection head outputs a "green light" status, the vehicle passage behavior output by the planning control head should be "pass," not "stop." Through this constraint, traffic rule violations can be detected and penalized.
[0071] In this embodiment, by introducing at least one compatibility constraint from the physical space constraints, causal relationship constraints, and application scenario business constraints of two associated task output heads in the multi-task joint model when executing the corresponding target task, spatial location conflicts, causal logic conflicts, and business rule conflicts between the target task results output by the two types of task output heads can be automatically identified. The severity of each type of conflict is then quantified, and a first self-consistent loss value is calculated. Subsequently, this first self-consistent loss value is fused with the basic loss value of the multi-task joint model to jointly update and optimize the model parameters of the multi-task joint model.
[0072] Step S330: Update the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model.
[0073] The basic loss value can be obtained based on the deviation between the target task result set and the corresponding ground truth result set of the sample data. For example, for a vehicle detection head, the basic loss value can be calculated based on the positional deviation between the vehicle bounding box parameters and the ground truth bounding box (such as intersection-over-union loss, center point offset loss, etc.); for a curb detection head, the basic loss value can be calculated based on the positional deviation between the curb point sequence and the ground truth curb points (such as point distance loss, curve fitting loss, etc.). The basic loss value is used to constrain the task accuracy of each task output head, so that the target task result output by the model is as close as possible to the ground truth.
[0074] In one embodiment, step S330, which updates the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model, may further include: calculating the total loss value of the multi-task joint model based on the first self-consistent loss value and its corresponding weight coefficients, as well as the basic loss value of the multi-task joint model; and updating the model parameters of the multi-task joint model based on the total loss value.
[0075] The weighting coefficient can be used to adjust the proportion of the first self-consistent loss value in the total loss value. For example, the weighting coefficient of the first self-consistent loss value can be preset. , The value can be between 0 and 1 (such as 0.1, 0.3, 0.5, etc.), and the specific value can be set according to the actual application scenario and conflict tolerance.
[0076] The total loss value can be calculated by weighted summation of the first self-consistent loss value and the basic loss value. For example, the total loss value... ,in, This is the basic loss value. The first self-consistent loss value, The weighting coefficients corresponding to the first self-consistent loss value.
[0077] In some examples, when the multi-task joint model includes multiple associated task output head pairs, and each task output head pair corresponds to a first self-consistent loss value, the total loss value can be expressed as: .in, Let be the first self-consistent loss value corresponding to the i-th pair of task output headers. This represents the weight coefficient corresponding to the i-th pair of task output headers. Different pairs of task output headers can be set with the same or different weight coefficients.
[0078] In some examples, the first self-consistent loss value is fused with the basic loss value of the multi-task joint model to construct the total loss. The gradient is backpropagated based on the total loss to jointly update the model parameters of the multi-task joint model. This allows the model to optimize the accuracy of each task while actively learning the compatibility between different task output heads, thereby reducing or eliminating spatial location conflicts, causal logic conflicts, and business rule conflicts across output heads.
[0079] In some examples, after completing the current model parameter update, the process can return to step S310 to obtain the target task result set output by the updated multi-task joint model for the new round of sample data, recalculate the first self-consistent loss value and the total loss value, and continue to update the model parameters until the preset iteration stopping condition is met. The preset iteration stopping condition may include the total loss value being less than a preset threshold and / or reaching the maximum number of iterations.
[0080] It should be noted that resolving the aforementioned conflicts mainly relies on manually labeled joint training data or algorithm engineers tuning parameters based on experience, but this suffers from problems such as separation of evaluation and training, long iteration cycles, and high costs. Therefore, how to efficiently detect output conflicts between different output heads in a multi-task joint model and transform the detection results into training optimization signals is a technical problem that urgently needs to be solved in this field. Through the embodiments provided in this disclosure, by introducing compatibility constraints between two related task output heads in the multi-task joint model when executing their respective target tasks, the output conflicts between the target task results output by the two task output heads are automatically detected and quantified into a first self-consistent loss value. The first self-consistent loss value is combined with the basic loss value of the multi-task joint model to jointly update the model parameters of the multi-task joint model, realizing closed-loop collaboration between conflict detection and model training. This eliminates the need to rely on manually labeled data or experience-based parameter tuning, significantly improving model iteration efficiency and reducing costs.
[0081] In some embodiments of this disclosure, step S320, which determines the first self-consistent loss value between the target task results output by two associated task output heads when executing their respective target tasks based on compatibility constraints among multiple task output heads, may further include the following steps S321 to S323: Step S321: Obtain the first task result output by the first task output head and the second task result output by the second task output head, which are associated among multiple task output heads.
[0082] For example, continuing with the multi-task joint model as the bird's-eye view BEV joint training model, the vehicle bounding box parameters output by the vehicle detection head can be obtained as the first task result, and the continuous roadside point sequence output by the roadside detection head can be obtained as the second task result.
[0083] For example, continuing with the multi-task joint model as the parking joint training model, we can obtain the parking space polygon parameters output by the parking space recognition head as the first task result, and obtain the binary mask output by the passable area detection head as the second task result.
[0084] Step S322: Generate a self-consistent evaluation index based on the compatibility constraints of the first task output head and the second task output head when executing their respective target tasks, the results of the first task and the second task.
[0085] Among them, the self-consistency evaluation index can be used to determine the degree of conflict between the results of the first task and the results of the second task.
[0086] In some examples, the compatibility constraint may include a compatibility feature and a feature evaluation threshold corresponding to the compatibility feature, where the compatibility feature is used to characterize the degree of conflict between target task results output by two task output heads respectively. In this example, determining the first self-consistency loss value between the first task result output by the first task output head and the second task result output by the second task output head according to the self-consistency evaluation index may be specifically implemented in the following manner: determining feature values of the compatibility feature corresponding to the first task result and the second task result; generating the self-consistency evaluation index for the compatibility feature corresponding to the first task result and the second task result according to the feature values of the compatibility feature and the feature evaluation threshold corresponding to the compatibility feature.
[0087] Continuing to take the multi-task joint model as the bird's eye view BEV joint training model as an example, the compatibility constraint may be a physical space constraint between the vehicle detection head and the curb detection head, the compatibility feature may be the minimum distance between the vehicle bounding box and the curb point sequence, and the feature evaluation threshold may be a preset safety distance (e.g., 0.5 meters). Specifically, the vehicle bounding box B output by the vehicle detection head can be obtained, and the vehicle bounding box B can be represented as (x0, y0, w, h, θ), where (x0, y0) is the center coordinate of the bounding box, w is the width, h is the height, and θ is the rotation angle; the curb point sequence R1(x1, y1), R2(x2, y2), ..., R output by the curb detection head is obtained n (x n , y n ), and convert the curb point sequence into a continuous set of curb line segments S={R1R2, R2R3, …, R n-1 R n}. Then, calculate the minimum distance d between the vehicle bounding box B and each line segment in the line segment set S i (which can be calculated by the shortest distance algorithm between a rectangle and a line segment), and take all d i the minimum value d_min among them as the feature value of the compatibility feature, and d_min is used to characterize the minimum separation distance between the vehicle and the curb. Next, the self-consistency evaluation index is generated according to the relationship between the feature value d_min and the feature evaluation threshold d0. Specifically, the self-consistency evaluation index can be expressed as: when d_min≥d0, the self-consistency evaluation index=0 (indicating no collision risk); when d_min<d0, the self-consistency evaluation index=d0-d_min, and the larger the difference, the closer the vehicle is to the curb, and the higher the collision risk.
[0088] Continuing to take the multi-task joint model as an example of the parking joint training model, the compatibility constraint may be the physical space constraint between the parking space recognition head and the passable area detection head, the compatibility feature may be the area proportion of the impassable area within the parking space polygon area, and the feature evaluation threshold may be a preset intrusion threshold (e.g., 5%). Specifically, the parking space polygon area P output by the parking space recognition head can be obtained, and the parking space polygon area P can be defined by vertices P1(x1, y1), P2(x2, y2), …, P k (x k , y k ) to characterize the contour shape and position of the parking space; obtain the binary mask M output by the passable area detection head, wherein in the binary mask M, a pixel value of 0 represents a passable area, and a pixel value of 1 represents an impassable area (such as obstacles, curbs, occupied parking spaces, etc.). Then, map the parking space polygon area P to the pixel coordinate system of the binary mask M to obtain the coverage area A_p of the parking space on the mask; count the number of pixels N_invade with pixel value 1 (i.e., impassable area) in the coverage area A_p, and the total number of pixels N_total of the coverage area A_p; calculate the intrusion ratio r=N_invade / N_total as the feature value of the compatibility feature, and the intrusion ratio r is used to characterize the severity of the intrusion of the impassable area into the parking space. Then, a self-consistency evaluation index is generated according to the relationship between the feature value r and the feature evaluation threshold r0. Specifically, the self-consistency evaluation index can be expressed as: when r≤r0, the self-consistency evaluation index=0, indicating that no passable area intrudes into the parking space or the intrusion degree is acceptable; when r>r0, the self-consistency evaluation index=r-r0, the larger the difference is, the more serious the intrusion of the impassable area into the parking space is, and the lower the passable safety of the parking space area is.
[0089] Step S323, determining the first self-consistency loss value between the first task result output by the first task output head and the second task result output by the second task output head according to the self-consistency evaluation index.
[0090] In this embodiment, the self-consistency evaluation index value can be directly determined as the first self-consistency loss value, so that the loss value is positively correlated with the conflict degree; the self-consistency evaluation index can also be subjected to non-linear transformation (such as square transformation) and then used as the first self-consistency loss value to strengthen the penalty for severe conflicts.
[0091] Continuing to take the multi-task joint model as an example of the bird's-eye view BEV joint training model, the first self-consistency loss value is determined according to the self-consistency evaluation index . Specifically, the first self-consistency loss value can be expressed by the following calculation rule: when d_min≥d0, f4=0; when d_min<d0, f4=d0-d_min. It should be noted that, if it is necessary to strengthen the penalty for severe collision risks, the square loss form can also be adopted, that is =(d0-d_min)². This first self-consistent loss value can be... Basic loss value of multi-task joint model Combined to construct the total loss value: .in, The base loss value can be calculated based on the deviation between the output of the vehicle detection head and the true value, and the deviation between the output of the curb detection head and the true value. The weighting coefficient for the first self-consistent loss value is used to adjust the proportion of compatibility constraints in the total loss.
[0092] Continuing with the example of a multi-task joint model for parking joint training, the first self-consistent loss value is determined based on this self-consistent evaluation index. Specifically, the first self-consistent loss value The calculation rule can be expressed as: when r ≤ r0, =0; when r>r0 =r-r0. It should be noted that if it is necessary to strengthen the punishment for serious intrusions, the squared loss form can also be used, i.e. =(r-r0)². This first self-consistent loss value can be... Basic loss value of multi-task joint model Combined, construct the total loss value: .in, The base loss value can be calculated based on the deviation between the output of the parking space recognition head and the ground truth (such as the deviation of the corner points of the parking space polygon) and the deviation between the output of the passable area detection head and the ground truth (such as the cross-entropy loss of semantic segmentation). λ is the weight coefficient of the first self-consistent loss value, used to adjust the proportion of compatibility constraints in the total loss. By minimizing the total loss through gradient descent, the model optimizes the accuracy of each task while maintaining spatial consistency between the parking space area and the passable area, preventing impassable areas from encroaching on parking spaces.
[0093] Through embodiments of this disclosure, by introducing compatibility constraints and compatibility features, the degree of conflict between the outputs of different task output heads can be automatically calculated and quantified into a first self-consistent loss value. This loss value is then combined with the basic loss value to update the model parameters. This method does not rely on manually labeled joint training data to achieve automatic detection and targeted optimization of cross-output head conflicts, effectively solving compatibility issues such as vehicle-curb collisions and intrusion into parking spaces in impassable areas, and significantly improving the training efficiency and output reliability of multi-task joint models.
[0094] In some embodiments of this disclosure, the update method of the multi-task joint model of this disclosure may further include: determining a second self-consistent loss value of the target task result output by each task output head based on the single-task self-consistency constraint of each task output head when executing the corresponding target task.
[0095] Among them, single-task self-consistency constraints can include internal logical consistency constraints, physical world rationality constraints, and / or business scenario constraints.
[0096] The inherent logical consistency constraint is used to constrain the structural continuity and order of the target task results output by a single task output head. For example, taking a lane detection head as an example, its output lane lines are an ordered sequence of points. The inherent logical consistency constraint requires that the arrangement of points conforms to the continuous extension characteristics of lane lines, and there should be no abrupt bends or discontinuities. If the included angle between the vectors of three consecutive points is too large (e.g., greater than 90°), it can be determined that there is a violation of the inherent logical consistency constraint.
[0097] Physical world rationality constraints are used to ensure that the target task results output by a single task output head conform to physical laws or real-world rules. For example, taking a left and right boundary line detection head as an example, its output left and right boundary lines are physical entities separating driving areas. Physical world rationality constraints require that they be parallel and spaced within a reasonable range, and should not intersect. If an intersection point is detected between the left and right boundary lines, it can be determined that there is a violation of the physical world rationality constraints.
[0098] Business scenario constraints are used to ensure that the target task results output by a single task output head conform to preset business rules. For example, taking a traffic sign recognition head, business scenario constraints require that the recognition results of adjacent frames maintain semantic consistency and should not contain contradictions (e.g., the previous frame recognizes a speed limit of 60 km / h, and the next frame recognizes a no-entry sign). If the semantic similarity of the recognition results of adjacent frames is detected to be lower than a preset similarity threshold, it can be determined that there is a violation of the business scenario constraints.
[0099] In this embodiment, step S330, which updates the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model, may further include: updating the model parameters of the multi-task joint model based on the second self-consistent loss value, the first self-consistent loss value, and the basic loss value of the multi-task joint model.
[0100] Continuing with the example of the multi-task joint model for the bird's-eye view BEV joint training model, this model can introduce not only the first self-consistent loss value between the vehicle detection head and the curb detection head (used to constrain the vehicle from colliding with the curb), but also the second self-consistent loss value of each task output head itself. For example, for the curb detection head, its output curb point sequence should be a continuous and smooth curve, without abrupt bends or discontinuities.
[0101] Continuing with the multi-task joint model as an example of the parking joint training model, this model can introduce not only the first self-consistent loss value between the parking space recognition head and the passable area detection head (used to constrain non-passable areas from intruding into parking spaces), but also the second self-consistent loss value of each task output head itself. For example, for the parking space recognition head, the output parking space polygon should be a closed convex quadrilateral, and there should be no abnormalities such as edge intersections or self-intersections.
[0102] Through embodiments of this disclosure, by introducing compatibility constraints between different task output heads, output conflicts across output heads are automatically detected and quantified to obtain a first self-consistent loss value. Simultaneously, by introducing single-task self-consistency constraints for each task output head, anomalies in the output of a single task output head are quantified to obtain a second self-consistent loss value. Combining these self-consistent loss values with the basic loss value to update the model parameters allows the model to optimize task accuracy while simultaneously considering compatibility between task output heads and self-consistency within task output heads, effectively improving the reliability and consistency of the model output.
[0103] In some embodiments of this disclosure, the method for updating the multi-task joint model may further include: obtaining the updated multi-task joint model; obtaining the set of verification results output by the updated multi-task joint model for verification data; calculating the feature values of the compatibility features based on the verification results output by the two associated task output heads in the set of verification results; adjusting the weight coefficient of the corresponding first self-consistent loss value when the feature value exceeds the feature evaluation threshold, and re-executing the above steps to calculate the total loss value of the multi-task joint model based on the first self-consistent loss value, its corresponding weight coefficient, and the basic loss value of the multi-task joint model.
[0104] Specifically, after the model is updated, the validation results output by the updated model for the validation data are obtained, and the feature values of the validation results output by the two related task output heads for the compatibility feature are calculated. If the feature value exceeds the preset feature evaluation threshold, it indicates that the model still has cross-task output head conflicts. In this case, the weight coefficient of the first self-consistent loss value is adjusted, and the total loss value is recalculated to continue training the model until the model output meets the compatibility constraint requirements.
[0105] Through the embodiments of this disclosure, a closed-loop iteration of "training-verification-adjustment-retraining" is achieved by introducing a verification feedback mechanism. This mechanism can automatically detect whether compatibility issues still exist after model updates and dynamically adjust the self-consistent loss weights based on the detection results. The model can be continuously optimized without manual intervention, significantly improving the model's iteration efficiency and final output quality, and ensuring that the model meets the expected compatibility requirements before actual deployment.
[0106] Figure 4 This is a flowchart illustrating an update method for a multi-task collaborative model provided in this disclosure. The update method for this multi-task collaborative model can be... Figure 1 The illustrated mobile device and / or server execute. For example... Figure 4 As shown, the control method of this embodiment may include the following steps S410 to S460.
[0107] Step S410: Obtain the target task result set output by the multi-task joint model for the sample data.
[0108] Among them, the sample data can be data generated based on scene data collected by mobile devices, the multi-task joint model can include multiple task output heads, different task output heads correspond to different target tasks, and the target task result set can include the target task results obtained by each task output head executing the corresponding target task.
[0109] Step S420: Obtain the first task result output by the first task output head and the second task result output by the second task output head, which are associated among multiple task output heads.
[0110] Step S430: Generate a self-consistent evaluation index based on the compatibility constraints of the first task output head and the second task output head when executing their respective target tasks, the results of the first task and the second task.
[0111] Step S440: Determine the first self-consistency loss value between the first task result output by the first task output head and the second task result output by the second task output head based on the self-consistency evaluation index.
[0112] Step S450: Based on the single-task self-consistency constraint of each task output head when executing the corresponding target task, determine the second self-consistency loss value of the target task result output by each task output head.
[0113] Step S460: Update the model parameters of the multi-task joint model based on each second self-consistent loss value, each first self-consistent loss value, and the basic loss value of the multi-task joint model.
[0114] Using the above method, output conflicts between different task output heads in a multi-task joint model can be automatically detected and quantified as a first self-consistent loss value. Simultaneously, output anomalies within a single task output head can be detected and quantified as a second self-consistent loss value. These self-consistent loss values are then combined with the basic loss value to update the model parameters. This method does not rely on manually labeled joint training data and can simultaneously optimize across output head conflicts and single output head anomalies, effectively improving the model's update efficiency and output reliability.
[0115] Figure 5 This is a flowchart illustrating a control method provided in an embodiment of this disclosure. The control method can be... Figure 1 This can be executed by the shown mobile device and / or server. It can also be executed by any electronic device. For example... Figure 5 As shown, the method may include steps S510 to S530.
[0116] Step S510: Acquire real-time input data collected by the mobile device.
[0117] The real-time input data can be environmental data collected in real time by sensors on mobile devices (such as vehicles, robots, etc.) during operation. For example, this real-time input data may include image data collected by an image sensor, point cloud data collected by a LiDAR, obstacle distance data collected by an ultrasonic radar, etc. This real-time input data is used to feed into a pre-trained multi-task joint model to obtain the real-time task results of each task output head.
[0118] Step S520: Input the real-time input data into the multi-task joint model to obtain the real-time task results output by each task output head in the multi-task joint model.
[0119] The multi-task joint model can be trained using the training method for multi-task joint models from any of the above method embodiments. The multi-task joint model is stored on a mobile device or a server corresponding to the mobile device.
[0120] For example, the multi-task joint model can be a bird's-eye view BEV joint training model. After real-time input data is input into the model, the vehicle detection head can output real-time vehicle bounding box parameters, and the roadside detection head can output real-time roadside point sequences.
[0121] For example, this multi-task joint model can also be a parking joint training model. After real-time input data is fed into this model, the parking space recognition head can output real-time parking space polygon parameters, and the passable area detection head can output a real-time binary mask. The above real-time task results can be used to assist mobile devices in environmental perception and motion control.
[0122] Step S530: Motion control is performed on the mobile device based on the real-time task results output by each task output head in the multi-task joint model.
[0123] For example, for the BEV joint training model, the position of surrounding vehicles can be perceived based on the real-time vehicle bounding box output by the vehicle detection head, and the road boundary can be perceived based on the real-time road edge point sequence output by the road edge detection head. Based on the above perception results, the driving path of the mobile device can be planned, and the mobile device can be controlled to avoid surrounding vehicles and drive along the road boundary.
[0124] For example, for a parking joint training model, the location and range of available parking spaces can be determined based on the real-time parking space polygon output by the parking space recognition head, the distribution of obstacles can be identified based on the real-time binary mask output by the passable area detection head, and a parking path can be planned based on the above perception results to control the mobile device to safely park in the parking space.
[0125] By using the above method, the trained multi-task joint model is used to infer real-time input data, and the motion control of the mobile device is performed based on the real-time task results of each task output head, thereby improving the perception capability and control accuracy of the mobile device in complex environments.
[0126] Figure 6 A schematic diagram of the structure of an electronic device is provided in the disclosed embodiments. For example... Figure 6 As shown, the electronic device 1000 may include a memory 1010 and a processor 1020. The memory 1010 may be used to store computer instructions, and the processor 1020 may be used to retrieve computer instructions from the memory 1010 to execute all or part of the steps of any of the methods in the foregoing embodiments of this disclosure. The processor may be one or more processors, which may execute instructions individually or jointly. Similarly, the memory may be one or more memories, which may store the aforementioned computer instructions individually or jointly.
[0127] In some examples, the electronic device can be Figure 1 The electronic device is a server and / or a mobile device. In other examples, the electronic device can also be any electronic device, such as a controller for a mobile device.
[0128] This disclosure also provides a mobile device that may include a memory and a processor. The memory may be used to store computer instructions, and the processor may be used to retrieve the computer instructions from the memory to perform all or part of the steps of any of the methods in the foregoing embodiments of this disclosure. The processor may be one or more processors, which may execute the instructions individually or jointly. Similarly, the memory may be one or more memories, which may store the aforementioned computer instructions individually or jointly.
[0129] The mobile device provided in this embodiment can be... Figure 1 or Figure 2 The mobile device shown is, in some examples, a vehicle, which may be an electric vehicle, a hybrid vehicle, a fuel cell vehicle, or another type of vehicle. The vehicle may be an autonomous vehicle or a non-autonomous vehicle.
[0130] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods in the foregoing embodiments of this disclosure. Optionally, the computer-readable storage medium may be a non-transitory storage medium, but is not limited thereto, and may also be a temporary storage medium.
[0131] This disclosure also provides a chip that may include a processing unit, which can be used to execute all or part of the steps of any of the methods in the foregoing embodiments of this disclosure. The chip may be in the form of an Application-Specific Integrated Circuit (ASIC), a System-on-Chip (SOC), a Field-Programmable Gate Array (FPGA), etc., and this embodiment is not limited to this. Optionally, the chip may further include a storage unit, which can be used to store computer instructions. The processing unit can be used to retrieve the computer instructions from the storage unit to execute all or part of the steps of any of the methods in the foregoing embodiments of this disclosure.
[0132] This disclosure also provides a computer program product that may include a computer program that, when executed by a processor, can implement any of the methods described in the foregoing embodiments of this disclosure.
[0133] This disclosure may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement any of the methods in the foregoing embodiments of this disclosure.
[0134] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0135] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0136] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0137] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0138] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0139] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It should be noted that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are all equivalent.
[0141] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of this disclosure is defined by the appended claims.
Claims
1. An update method for a multi-task joint model, characterized in that, The method includes: Obtain the target task result set output by the multi-task joint model for the sample data; wherein, the sample data is data generated based on scene data collected by a mobile device, the multi-task joint model includes multiple task output heads, different task output heads correspond to different target tasks, and the target task result set includes the target task results obtained by each of the task output heads performing the corresponding target task; Based on the compatibility constraints of two associated task output heads among the plurality of task output heads when executing their respective target tasks, a first self-consistency loss value is determined between the target task results output by the two task output heads. The model parameters of the multi-task joint model are updated based on the first self-consistent loss value and the basic loss value of the multi-task joint model; wherein the basic loss value is obtained based on the deviation between the target task result set and the true result set corresponding to the sample data.
2. The method according to claim 1, characterized in that, The step of determining the first self-consistency loss value between the target task results output by two associated task output heads when executing their respective target tasks, based on the compatibility constraints of these two task output heads, includes: Obtain the first task result output by the first task output head and the second task result output by the second task output head associated with the plurality of task output heads; Based on the compatibility constraints of the first task output head and the second task output head when executing their respective target tasks, the first task result and the second task result, a self-consistent evaluation index is generated; wherein, the self-consistent evaluation index is used to determine the degree of conflict between the first task result and the second task result; The first self-consistency loss value between the first task result output by the first task output head and the second task result output by the second task output head is determined based on the self-consistency evaluation index.
3. The method according to claim 2, characterized in that, The compatibility constraint includes compatibility features and feature evaluation thresholds corresponding to the compatibility features. The compatibility features are used to characterize the degree of conflict between the target task results output by the two task output heads respectively. The step of generating a self-consistent evaluation index based on the compatibility constraints of the first task output header and the second task output header when executing their respective target tasks, the first task result, and the second task result includes: Determine the feature values of the first task result and the second task result for the compatibility feature; Based on the feature value of the compatibility feature and the feature evaluation threshold corresponding to the compatibility feature, a self-consistent evaluation index for the first task result and the second task result with respect to the compatibility feature is generated.
4. The method according to claim 3, characterized in that, The step of updating the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model includes: The total loss value of the multi-task joint model is calculated based on the first self-consistent loss value and its corresponding weight coefficient, as well as the basic loss value of the multi-task joint model. The model parameters of the multi-task joint model are updated based on the total loss value.
5. The method according to claim 1, characterized in that, The compatibility constraints include at least one of the following: physical space constraints, causal relationship constraints, and application scenario business constraints; The physical space constraint is used to constrain the spatial position matching relationship between the target task results output by different task output heads; The causal association constraint is used to constrain the causal logical consistency relationship between the target task results output by different task output heads; The application scenario business constraints are used to constrain the matching relationship between the target task results output by different task output heads and the preset business rules.
6. The method according to claim 4, characterized in that, The method further includes: Obtain the updated multi-task joint model; Obtain the updated set of verification results output by the multi-task joint model for the verification data; The feature value of the compatibility feature is based on the verification results output by the two associated task output headers in the verification result set; If the feature value exceeds the feature evaluation threshold, adjust the weight coefficient of the corresponding first self-consistent loss value, and re-execute the step of calculating the total loss value of the multi-task joint model based on the first self-consistent loss value, its corresponding weight coefficient, and the basic loss value of the multi-task joint model.
7. The method according to claim 1, characterized in that, The method further includes: Based on the single-task self-consistency constraint of each of the plurality of task output heads when executing the corresponding target task, a second self-consistency loss value of the target task result output by each of the task output heads is determined; The step of updating the model parameters of the multi-task joint model based on the first self-consistent loss value and the basic loss value of the multi-task joint model includes: The model parameters of the multi-task joint model are updated based on the second self-consistent loss value, the first self-consistent loss value, and the basic loss value of the multi-task joint model.
8. A control method, characterized in that, The method includes: Acquire real-time input data collected by mobile devices; The real-time input data is input into the multi-task joint model to obtain the real-time task results output by each task output head in the multi-task joint model; wherein, the multi-task joint model is trained by the training method of the multi-task joint model according to any one of claims 1 to 7; Motion control is performed on the mobile device based on the real-time task results output by each task output head in the multi-task joint model.
9. An electronic device, characterized in that, The method includes a memory and a processor, the memory being used to store computer instructions, and the processor being used to retrieve the computer instructions from the memory to perform the method of any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 8.