Mobile Edge Intelligent Inference Method and System Supporting Multimodal Input
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]然而,现有的对人员需求数据进行识别处理的技术在数据的智能推理方面存在明显缺陷,例如,现有方法大多对单一模态的数据进行分析推理,往往无法对人员的多个模态的数据进行综合解析,例如,在实际会议场景中,面对会议过程中参会人员的语音数据、视频数据、文字数据等多种类型的数据时,现有技术可能难以将这些不同模态的数据进行分析处理,可能无法全面理解参会人员通过不同方式传达的需求信号,从而可能无法精准识别参会人员在会议过程中的各种需求,导致不能及时对参会人员的需求进行处理
1、本发明通过对参会人员对应的多种模态的数据进行分析处理,可以更加全面、更加准确的对参会人员在会议过程中的需求信息进行识别,从而可以更加精准的为参会人员提供相适配的服务,提升参会人员的体验。
Smart Images

Figure CN120997879B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to data processing technology, and more particularly to a mobile edge intelligent reasoning method and system that supports multimodal input. Background Technology
[0002] With the widespread adoption of mobile devices and the rapid development of artificial intelligence technology, mobile edge intelligent reasoning is playing an increasingly crucial role in many fields. For example, in meeting scenarios, edge-side intelligent reasoning can accurately identify the various needs of participants, improve meeting quality, and provide technical support for the intelligent development of meetings.
[0003] However, existing technologies for identifying and processing personnel needs data have significant shortcomings in intelligent reasoning. For example, most existing methods analyze and reason about data in a single modality, and often cannot comprehensively analyze data from multiple modalities. For instance, in actual meeting scenarios, when faced with various types of data such as voice data, video data, and text data from participants, existing technologies may struggle to analyze and process these different modalities. They may not be able to fully understand the needs signals conveyed by participants in different ways, and thus may not be able to accurately identify the various needs of participants during the meeting, resulting in a failure to address the needs of participants in a timely manner.
[0004] Therefore, how to analyze and process multimodal data to more comprehensively and accurately identify the needs of participants has become an urgent problem to be solved. Summary of the Invention
[0005] This invention provides a mobile edge intelligent inference method and system that supports multimodal input, which can analyze and process multimodal data to more comprehensively and accurately identify the needs of participants.
[0006] A first aspect of the present invention provides a mobile edge intelligent inference method supporting multimodal input, comprising: In response to the pre-completion information, the system receives the target image collected by the monitoring equipment for the target area, identifies and stitches together the static planning area and dynamic stitching area of each target element in the target image to obtain the target recognition area of the target element. Based on the target amplitude of the target element in the target recognition area, the corresponding target recognition area is used as the trigger recognition area, and the multimodal data of the trigger recognition area is retrieved. The multimodal data includes input data and output data. The input data is parsed to obtain active adjustment information corresponding to the input data. Based on the change recognition strategy, the dynamic changes of the main element and the accompanying element in the output data are identified to obtain the trigger adjustment information of the output data. The active adjustment information and the triggered adjustment information are sent to the adjustment device of the target element.
[0007] Optionally, in one possible implementation of the first aspect, the step of identifying and stitching together the static planning region and dynamic stitching region of each target element in the target image to obtain the target recognition region of the target element includes: Identify the positional contours of statically positioned elements and the placement contours of statically placed elements in the target image; Determine the position contour that is closest to each placement side length in the placement contour as the processing contour for each placement side length; The fixed side length and extended side length in the processing contour are processed to obtain the static planning area of the target element; Based on the placement side length, the corresponding static planning area is mirrored to obtain the copied planning area; The copy planning area is dynamically filled in according to the accompanying outline of the accompanying placement element in the placement outline to obtain the dynamic splicing area of the target element. The static planning area and the dynamic splicing area are spliced together to obtain the target recognition area of the target element.
[0008] Optionally, in one possible implementation of the first aspect, processing the fixed side length and extended side length in the processing contour to obtain the static planning region of the target element includes: The side length parallel to the corresponding placement side length in the processing contour is taken as the fixed side length, and the remaining side length is taken as the extended side length. Delete the fixed side length closest to the corresponding placement side length, and extend the extended side length to one side of the corresponding placement side length until it intersects with the placement side length, thus obtaining the static planning area of the target element.
[0009] Optionally, in one possible implementation of the first aspect, the step of dynamically filling the copy planning area according to the accompanying contour of the accompanying placement element in the placement contour to obtain the dynamic splicing area of the target element includes: Identify the accompanying contours of the accompanying placement elements in the placement contour, take the accompanying contours that intersect with the copy planning area as the merged contours, and take the remaining accompanying contours as the contours to be selected. The first combined contour is obtained by the union of the merged contour and the corresponding replicated planning area; The target image is processed by coordinate conversion to obtain the straight-line distance between the center point of the contour to be selected and the first combined contour. The contour to be selected whose straight-line distance is less than the preset filtering distance is selected as the corresponding first combined contour. Obtain the extreme coordinate values of the selected contour and the first combined contour, construct a reconstructed rectangle based on the extreme coordinate values, and use the reconstructed rectangle as the dynamic splicing area of the target element.
[0010] Optionally, in one possible implementation of the first aspect, the step of using the corresponding target recognition region as the trigger recognition region based on the target amplitude of the target element in the target recognition region includes: Identify the target amplitude of the target element in the target identification area, and when the target amplitude is greater than a preset amplitude, use the corresponding target identification area as the trigger identification area.
[0011] Optionally, in one possible implementation of the first aspect, the step of identifying dynamic changes in the main elements and accompanying elements in the output data based on the change recognition strategy to obtain trigger adjustment information for the output data includes: Acquire output data, which includes video data and audio data; Based on the dynamic changes of the main elements and accompanying elements in the audio data using the audio recognition strategy, the triggering and adjustment information of the audio data is obtained. Based on the video recognition strategy, the dynamic changes of the main elements and accompanying elements in the video data are analyzed to obtain the trigger adjustment information of the video data.
[0012] Optionally, in one possible implementation of the first aspect, obtaining trigger adjustment information for the audio data based on the dynamic changes of the main elements and accompanying elements in the audio data using an audio recognition strategy includes: Based on timbre recognition, the main sound wave in the audio data is taken as the main element, and the remaining sound waves are taken as accompanying elements; Extract the main amplitude of the main element in the audio data, and when the main amplitude is greater than the preset main amplitude, take the corresponding sound wave as the main change sound wave, and obtain the main duration of each main change sound wave; When the duration of the main body is determined to be less than the preset duration, the corresponding change in the main body sound wave is taken as the instantaneous main body sound wave, and the preset transmission information is retrieved based on the instantaneous main body sound wave; When the duration of the main body is determined to be greater than the preset duration, the corresponding change sound wave of the main body is taken as the continuous main body sound wave, the text information corresponding to the continuous main body sound wave is extracted, and annotation adjustment information is generated based on the text information; Extract the accompanying amplitude of the accompanying element in the audio data. When the accompanying amplitude is greater than the preset accompanying amplitude, the corresponding sound wave is taken as the accompanying change sound wave, and the accompanying duration of each accompanying change sound wave is obtained. When it is determined that the accompanying duration is less than the preset duration, the corresponding accompanying change sound wave is taken as the instantaneous accompanying sound wave, and the processing and adjustment information is retrieved based on the instantaneous accompanying sound wave. When the duration of the accompanying sound wave is determined to be longer than a preset duration, the corresponding accompanying sound wave is taken as a continuous accompanying sound wave, and a cancellation sound wave opposite to the continuous accompanying sound wave is generated. The cancellation sound wave is used as cancellation adjustment information.
[0013] Optionally, in one possible implementation of the first aspect, obtaining trigger adjustment information for the video data based on the dynamic changes of the main elements and accompanying elements in the video data according to the video recognition strategy includes: Identify target elements in video data as main elements, and accompanying elements as accompanying elements; Identify the changing state of the corresponding element parts of the main element in the video data, and generate instruction adjustment information based on the changing state; Obtain the display attributes of accompanying elements in the video data, including transparency and concealment attributes; When the display attribute is determined to be transparent, the corresponding accompanying element is taken as an identifiable element, the usage height of the identifiable element is identified, and when the usage height is less than the limit height, displacement adjustment information is generated. When the display attribute is determined to be a hidden attribute, the corresponding accompanying element is treated as an unrecognizable element. The usage tilt of the unrecognizable element is identified. If the usage tilt is greater than the limit tilt, adjustment information is generated.
[0014] Optionally, in one possible implementation of the first aspect, the step of identifying the change state of the element corresponding to the main element in the video data and generating instruction adjustment information based on the change state includes: Identify the changing state of the corresponding element parts of the main element in the video data, wherein the changing state includes independent changes and combined changes; When the change state is determined to be an independent change, the change range of the element part is identified. When the change range of the part is determined to be greater than the preset part range, the corresponding preset instruction information is retrieved based on the element part as instruction adjustment information. When the change state is determined to be a combination change, the corresponding element part is taken as the combination part. When the combination parts have an intersection, the preset instruction information corresponding to the corresponding combination part is taken as the instruction adjustment information.
[0015] A second aspect of the present invention provides a mobile edge intelligent inference system supporting multimodal input, comprising: The stitching module is used to respond to the pre-completion information, receive the target image collected by the monitoring device on the target area, identify the static planning area and dynamic stitching area of each target element in the target image and stitch them together to obtain the target recognition area of the target element; The retrieval module is used to retrieve multimodal data of the trigger recognition area based on the target amplitude of the target element in the target recognition area, and the corresponding target recognition area is used as the trigger recognition area. The multimodal data includes input data and output data. The parsing module is used to parse the input data to obtain active adjustment information corresponding to the input data, and to identify the dynamic changes of the main elements and accompanying elements in the output data based on the change recognition strategy to obtain the trigger adjustment information of the output data. The sending module is used to send the active adjustment information and the triggered adjustment information to the adjustment device of the target element.
[0016] A third aspect of the present invention provides an electronic device comprising: a memory, a processor, and a computer program, the computer program being stored in the memory, and the processor executing the computer program to perform the methods described in the first aspect of the present invention and various possible methods related to the first aspect.
[0017] The beneficial effects of this invention are as follows: 1. By analyzing and processing data from multiple modalities of participants, this invention can more comprehensively and accurately identify the needs of participants during the meeting, thereby providing more precise and tailored services and enhancing the participants' experience.
[0018] 2. When acquiring multimodal data corresponding to attendees, this invention can first accurately obtain the target element, i.e., the target identification area corresponding to the attendees, through region recognition and stitching operations. Through the target identification area, multimodal data corresponding to the target element can be collected more comprehensively and accurately. Specifically, the processing contour can be determined by recognizing the position contour of static location elements and the placement contour of static placement elements. The side length of the processing contour is reasonably processed to obtain a static planning area. Then, a dynamic stitching area is obtained through mirror copying and dynamic completion processing. The static planning area and the dynamic stitching area are stitched together to obtain the target identification area corresponding to the target element. This can accurately delineate an area containing attendees and their related items, providing a comprehensive and accurate data foundation for subsequent demand identification based on this area, thereby improving the accuracy and reliability of demand identification.
[0019] 3. After determining the target recognition area, this invention can use the target amplitude as a screening criterion to quickly locate the area that truly needs attention, i.e. the trigger recognition area, in the meeting data. It can accurately retrieve the multimodal data of the trigger recognition area, comprehensively capture the various needs of the participants during the meeting, reduce unnecessary data processing, and improve the efficiency and accuracy of needs analysis.
[0020] 4. When analyzing multimodal data, this invention can perform in-depth analysis of both input and output data. Through different strategies and algorithms, it can accurately obtain the active adjustment information corresponding to the input data and the trigger adjustment information corresponding to the output data, thereby comprehensively reflecting the needs of the target elements during the meeting. Specifically, active adjustment information can be obtained by parsing the input data, and trigger adjustment information can be obtained by analyzing the dynamic changes of the main elements and accompanying elements in the output data using a change recognition strategy. This enables the comprehensive capture of the needs signals of the participants, including the needs explicitly expressed through the input data and the potential needs expressed through the output data, thereby improving the accuracy and comprehensiveness of the identification of the participants' needs.
[0021] 5. When analyzing the output data, this invention can perform more refined analysis on video and audio data separately. In terms of audio data processing, the dynamic changes of the main elements and accompanying elements in the audio data are analyzed through audio recognition strategies. Based on the changes in sound characteristics, the needs of the participants can be accurately determined. In terms of video data processing, the changes in the actions, expressions, and accompanying elements of the main elements in the video data are analyzed through video recognition strategies. Corresponding instruction adjustment information is then generated. Through refined analysis of audio and video data, service instructions can be accurately generated based on the specific state of the participants, achieving precise response to the needs of the participants and improving their meeting experience. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating a mobile edge intelligent push method supporting multimodal input provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of a dynamic splicing area for determining a target element provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a mobile edge intelligent push system supporting multimodal input provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0025] See Figure 1 This is a flowchart illustrating a mobile edge intelligent inference method supporting multimodal input provided by an embodiment of the present invention. Figure 1 The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. Steps S1 to S4 are detailed as follows: S1, responding to the pre-completion information, receiving the target image collected by the monitoring device for the target area, identifying the static planning area and dynamic splicing area of each target element in the target image and splicing them together to obtain the target recognition area of the target element.
[0026] Among them, "pre-conference completion information" refers to the information corresponding to the completion of all participants' pre-conference preparations. For example, when it is detected that all participants have placed their personal items such as water cups and pens, the corresponding pre-conference completion information can be displayed. "Monitoring equipment" refers to the equipment in the meeting venue that monitors the personnel in the corresponding area in real time. For example, it can be a surveillance camera installed in the meeting room. "Target area" refers to the area where the meeting is held. For example, it can be the area corresponding to the meeting room. "Target image" refers to the image obtained by the monitoring equipment from the target area. "Target element" refers to the participants in the meeting. "Static planning area" refers to the area corresponding to the location of the target element. "Dynamic stitching area" refers to the area containing all items corresponding to the target element, such as water cups and pens. "Target recognition area" refers to the area that can recognize the needs of the target element during the meeting.
[0027] Understandably, both internal corporate project seminars and academic conferences aim to identify the diverse needs of participants in real time and accurately to facilitate effective information transmission and rapid decision-making. However, most existing technologies can only perform reasoning and analysis on single-modal data, often failing to integrate and analyze data from multiple modalities during a meeting. For example, they may struggle to process participants' voice, video, and text data, potentially hindering the accurate identification of their needs and consequently preventing the provision of tailored services. This solution analyzes and processes data from multiple modalities, enabling a more comprehensive and accurate identification of participants' needs during meetings. This allows for more precise provision of appropriate services, enhancing the participant experience.
[0028] In practical applications, when the monitoring equipment detects that all participants in the meeting have completed their preparations, such as placing water cups, pens, and other items, it can send corresponding pre-completion information. Upon receiving this information, the monitoring equipment can acquire target images of the target area where the meeting is being held, obtained through image capture of the target area by a surveillance camera in the meeting room. These target images contain multiple target elements. By analyzing the target images, the location of each target element—that is, the location of the participant—can be identified, and its corresponding area can be defined as the static planning area. Simultaneously, the areas where the items of the target elements are located can be identified and used as dynamic stitching areas. By stitching the static planning area and the dynamic stitching area together, the target recognition area corresponding to each target element can be obtained. The target recognition area is the basic area for subsequent demand identification, integrating information about the target element and its related items, and providing image-level data support for a comprehensive analysis of the demand for the target elements.
[0029] In some embodiments, step S1, "identifying and stitching together the static planning region and dynamic stitching region of each target element in the target image to obtain the target recognition region of the target element," includes the following steps: S11, identify the position contours of static position elements in the target image, and the placement contours of static placement elements.
[0030] Among them, static position elements refer to the chairs where the participants sit, position outlines refer to the outlines corresponding to static position elements, static placement elements can be conference tables, and placement outlines refer to the outlines corresponding to static placement elements.
[0031] Specifically, image recognition technology can be used to perform preliminary analysis of the target image. Static position elements refer to the chairs where the participants are sitting. The position and outline of the chairs can reflect the basic position information of the participants in the meeting space. Static placement elements can be the conference table. The placement outline of the conference table can define the main planar area of the meeting activity and is also closely related to the position of the participants. Through edge detection, feature extraction and other algorithms, the boundary lines of the chairs in the image can be identified to form the corresponding position outline. At the same time, the edge of the conference table can be accurately located to obtain its corresponding placement outline. This outline information is the basis for subsequent area division and processing, providing a key reference for accurately determining the area where the participants are located.
[0032] S12, determine the position contour that is closest to each placement side length in the placement contour as the processing contour of each placement side length.
[0033] After obtaining the placement outlines corresponding to static placement elements and the position outlines corresponding to static position elements, the spatial relationship between the two can be analyzed. For each side length of the placement outline (conference table outline), the distance between it and each position outline (chair outline) can be calculated. The position outline closest to that side length can be found and determined as the processing outline corresponding to that placement side length. This step can clarify the spatial correspondence between the conference table and the chair, providing a basis for subsequent processing of the area based on this relationship. This ensures that the subsequently generated area can closely surround the participants and their surrounding environment, making the divided area more in line with the activity range of the participants in the actual scene.
[0034] Among them, the placement side length refers to the side length corresponding to the placement contour, and the processing contour refers to the position contour that is closest to each placement side length.
[0035] S13, process the fixed side length and extended side length in the processing contour to obtain the static planning area of the target element.
[0036] After determining the processing outline, the side lengths of the processing outline can be further classified and processed. Fixed side lengths refer to relatively stable sides that reflect the inherent shape and position of the chair, while extended side lengths are sides that may need to be extended or adjusted based on factors such as their positional relationship with the conference table. Through specific algorithms and rules, the fixed side lengths maintain their original shape, while the extended side lengths are appropriately extended. For example, based on the relative position of the conference table and the chair, the side of the chair perpendicular to the conference table can be reasonably extended to better cover the range of activities that participants may be involved in during the meeting. After processing the fixed and extended side lengths, a more accurate area can be obtained, namely the static planning area of the target element (participants). This area can more accurately define the basic position range of participants in the meeting space.
[0037] Among them, fixed side length refers to processing relatively fixed sides in the contour, while extended side length refers to processing side lengths in the contour that can be extended.
[0038] Based on the above embodiments, step S13 can be implemented in the following ways: S131, the side length parallel to the corresponding placement side length in the processing contour is taken as the fixed side length, and the remaining side length is taken as the extended side length.
[0039] Specifically, each edge of the processing contour can be analyzed one by one to determine whether it is parallel to a certain edge of the conference table placement contour. If an edge is parallel to the conference table placement edge, then this edge can be identified as a fixed edge. For the remaining edge lengths in the processing contour that are not parallel to the conference table placement edge, they can be defined as extended edge lengths. Through this edge length classification method based on parallel relationships, the subsequent processing direction for different edge lengths can be clarified, preparing for the construction of an accurate static planning area.
[0040] S132, delete the fixed side length closest to the corresponding placement side length, and extend the extended side length to one side of the corresponding placement side length until it intersects with the placement side length, thus obtaining the static planning area of the target element.
[0041] Specifically, after determining the fixed side lengths, the fixed side closest to the side where the conference table is placed can be deleted. Since this closest fixed side restricts the expansion of the area towards the conference table, and the conference table is an area frequently touched and used by participants during the meeting, deleting this side creates space for subsequent extensions, allowing the defined area to better cover the participants' activity range around the conference table. Then, the extensions can be made by extending each side towards the side where the corresponding side of the conference table is placed. This extension process continues until the extensions meet the side of the conference table. The area is defined by extending the edges of the table and chairs until they intersect. This ensures that the defined area closely matches the spatial relationship between the table and chairs, fully covering the activity areas that participants may engage in during the meeting. Once all extended edges intersect with the table's placement edge, the resulting area becomes the static planning area for the target element (participants). This area precisely defines the basic position range of participants in the meeting space, taking into account not only the inherent shape and position of the chairs but also the relative positional relationship between the table and chairs. This provides an accurate positional basis for subsequent identification and analysis of participant needs based on this area.
[0042] S14, based on the placement side length, mirror the corresponding static planning area to obtain the copied planning area.
[0043] Specifically, using the side length of the conference table as a reference, the generated static planning area can be mirrored. In a real meeting scenario, by mirroring the table with its side as the axis of symmetry, the area where attendees might have items or engage in activities on the table can be simulated. This provides a basic framework for considering the areas where all items around the attendees are located, allowing the generated area to more comprehensively cover the placement of items and the range of activities that attendees might be involved in. The mirrored planning area refers to the area obtained by mirroring the static planning area according to its side length.
[0044] S15, dynamically fill in the copy planning area according to the accompanying outline of the accompanying placement element in the placement outline to obtain the dynamic splicing area of the target element.
[0045] In practical applications, accompanying elements placed on the conference table by attendees may be located on the boundary of the copy planning area or outside the copy planning area. Accompanying elements can be various items placed on the conference table or around the attendees, such as water cups, pens, notebooks, etc. Their accompanying contours refer to the boundary shapes of these items in the image. Based on the accompanying contours of these accompanying elements, the copy planning area can be further processed. For the parts of the copy planning area that are not covered by accompanying elements, dynamic filling and expansion operations can be used to fill them. For example, if there is a water cup outside the copy planning area, the copy planning area can be expanded according to the contour of the water cup to include the area where the water cup is located. After this dynamic filling process, a dynamic splicing area corresponding to the target element can be obtained. This area completely includes the areas where all the items corresponding to the attendees are located, providing richer image information for a comprehensive analysis of the attendees' needs.
[0046] Among them, accompanying elements refer to various items placed on the conference table or around the participants, such as water cups, pens, notebooks, etc., and accompanying outlines refer to the outlines corresponding to the accompanying elements.
[0047] Based on the above embodiments, step S15 can be implemented in the following ways: S151, identify the accompanying contours of the accompanying placement elements in the placement contour, take the accompanying contours that intersect with the copy planning area as the merged contours, and take the remaining accompanying contours as the selection contours.
[0048] Specifically, image recognition technology can be used to identify various accompanying elements (such as water cups, pens, notebooks, etc.) within the placement contour of the conference table, and extract the boundary shapes of these items in the target image, i.e., the accompanying contours. The spatial relationship between each extracted accompanying contour and the previously generated copy planning area is determined based on whether the two areas intersect. If an accompanying contour overlaps with the copy planning area in the image space, then the accompanying contour is marked as a merged contour. The items corresponding to these merged contours are already partially within the copy planning area and need to be fully incorporated into the area later. For those accompanying contours that do not intersect with the copy planning area, they are classified as contours to be selected. These items are completely outside the copy planning area and need to be further determined whether to include them in the expansion range.
[0049] Among them, the merged contour refers to the accompanying contour that intersects with the replicated planning area, and the contour to be selected refers to the accompanying contour that does not intersect with the replicated planning area.
[0050] S152, based on the union of the merged contour and the corresponding copied planning area, a first combined contour is obtained.
[0051] After determining the merge outlines, these merge outlines can be combined with the corresponding copy planning areas using a union operation. The purpose of the union operation is to completely integrate the areas where the accompanying placement elements were originally located within the copy planning area into a new area. In this way, the copy planning area and the item areas represented by the accompanying outlines that intersect with it can be merged, eliminating the boundary barriers between areas and forming a more complete area outline that includes some accompanying placement elements, namely the first combined outline. Compared with the original copy planning area, this first combined outline has been initially expanded in scope and is more in line with the actual distribution of items used by the participants, which can prepare for further expansion of the area to include more items.
[0052] The first combined contour refers to the contour obtained by combining the merged contour with the corresponding replicated planning area.
[0053] S153, the target image is processed by coordinate conversion to obtain the straight-line distance between the center point of the contour to be selected and the first combined contour, and the contour to be selected whose straight-line distance is less than the preset filtering distance is selected as the corresponding first combined contour.
[0054] To further determine whether to include accompanying elements outside the copy planning area in the dynamic stitching area, the target image can be processed into coordinates. By establishing an image coordinate system, each pixel in the image is assigned a coordinate value, thus accurately calculating the spatial distance between different areas. For each candidate contour (i.e., an accompanying contour that does not intersect with the copy planning area), the coordinates of its center point can be obtained. Then, the shortest straight-line distance from the center point to each point on the first combined contour is calculated. By comparing the magnitude of these straight-line distances with the preset filtering distance, it is determined whether to include the candidate contour in the expansion range. The preset filtering distance is a threshold pre-set based on the possible distribution range of items of attendees in the actual meeting scenario. If the straight-line distance from the center point of a candidate contour to the first combined contour is less than the preset filtering distance, it means that although the item is outside the copy planning area, it is close to the integrated area and is likely an item that attendees will use during the meeting. Therefore, it is selected as the contour. Through this distance-based filtering method, while ensuring the integrity of the area, it is possible to avoid over-expanding the area, ensuring that the dynamic stitching area can cover important item areas without including too many irrelevant areas, thus improving the accuracy and effectiveness of area construction.
[0055] Wherein, straight line distance refers to the vertical distance between the center point of the contour to be selected and the first combined contour; preset filter distance refers to a pre-set distance used to measure the magnitude of straight line distance; and selected contour refers to the contour to be selected whose straight line distance is less than the preset filter distance.
[0056] S154, obtain the coordinate extreme values of the selected contour and the first combined contour, construct a reconstruction rectangle based on the coordinate extreme values, and use the reconstruction rectangle as the dynamic splicing area of the target element.
[0057] After the selected contour is determined, the selected contour and the first combined contour can be analyzed and processed. First, the extreme values of these contours in the image coordinate system can be obtained, that is, the minimum and maximum values of the horizontal coordinate and the minimum and maximum values of the vertical coordinate among all contour points. These extreme values can determine a minimum rectangular range that can completely contain the selected contour and the first combined contour. Based on the extreme values, a reconstruction rectangle corresponding to the minimum rectangular range can be constructed. This reconstruction rectangle is the final dynamic splicing area of the target element, which can completely contain the area where all items corresponding to the participants are located.
[0058] See Figure 2 This is a schematic diagram illustrating a dynamic splicing area for determining a target element, as provided in an embodiment of the present invention. Figure 2 As shown, after obtaining the first combined contour, when the straight-line distance between the contour to be selected and the first combined contour on the placement contour is less than the preset filtering distance, such as... Figure 2The circular outline in the diagram can be used to determine the corresponding outline to be selected as the selected outline. Based on the extreme coordinate values of the selected outline and the first combined outline, the smallest rectangle containing the selected outline and the first combined outline can be constructed, i.e., the reconstructed rectangle, as shown below. Figure 2 As shown, the area corresponding to the reconstructed rectangle can be defined as the dynamic splicing area.
[0059] By constructing a reconstructed rectangle, it is easier to manage and analyze the area in a unified manner. Whether it is retrieving image data from the area or performing multimodal data fusion analysis based on the area, it is more convenient and efficient. The formation of the dynamic stitching area provides a comprehensive and well-organized image information foundation for the accurate identification of participants' needs based on the target recognition area. It can more accurately analyze the relationship between participants and surrounding objects, thereby better understanding the needs of participants during the meeting.
[0060] Among them, the coordinate extreme values include the maximum horizontal coordinate value, minimum horizontal coordinate value, maximum vertical coordinate value, and minimum vertical coordinate value corresponding to the selected contour and the first combined contour. The reconstructed rectangle refers to the rectangle constructed based on the coordinate extreme values.
[0061] Based on the above embodiments, when the straight-line distance between the selected contour and multiple first combined contours is less than a preset filtering distance, meaning that there may be an item between at least two participants, this item may be shared by multiple participants, such as two participants sharing a laptop. In this case, the area corresponding to the shared item can be separately defined, and then a marker line can be used to connect the separately defined area with the dynamic splicing area corresponding to each participant, indicating that multiple participants share the corresponding item. Specifically, the contour edge of the selected contour that is closest to each first combined contour can be determined as the selected contour edge. For example, for contour edges 1, 2, 3, and 4 of the contour to be selected, when contour edge 1 is very close to the first combined contour 1 and contour edge 3 is very close to the first combined contour 2, contour edge 1 and contour edge 3 can be determined as contour edges to be selected. Then, the center point of the contour edge to be selected can be connected to the corresponding first combined contour. For example, the center point corresponding to contour edge 1 can be connected to the first combined contour 1, and the center point corresponding to contour edge 3 can be connected to the first combined contour 3, indicating that the item corresponding to the contour to be selected is an item shared by the participants corresponding to the first combined contour 1 and the participants corresponding to the first combined contour 3.
[0062] S16, the static planning area and the dynamic splicing area are spliced together to obtain the target recognition area of the target element.
[0063] After obtaining the static planning area and the dynamic stitching area, these two areas can be integrated. For example, image fusion technology can be used to seamlessly stitch the static planning area (representing the location of attendees) and the dynamic stitching area (containing the area of attendees' belongings) together to form a complete area, namely the target recognition area of the target element. The target recognition area integrates the location information of attendees and the information of surrounding items, which can become the basic area for identifying and analyzing the needs of attendees during the meeting. Subsequently, based on this target recognition area, multimodal data can be further retrieved to conduct a more in-depth and comprehensive analysis and judgment of the needs of attendees, laying a solid foundation for providing adapted services.
[0064] The above implementation method can accurately locate the possible activity areas of participants, reduce the interference of irrelevant information on demand judgment, and improve the pertinence of demand identification.
[0065] S2, based on the target amplitude of the target element in the target recognition area, the corresponding target recognition area is used as the trigger recognition area, and the multimodal data of the trigger recognition area is retrieved. The multimodal data includes input data and output data.
[0066] Among them, target amplitude refers to the amplitude of the target element's action during the meeting, trigger recognition area refers to the target recognition area corresponding to the target element with a large target amplitude, and multimodal data refers to data that can be used to infer the demand information of the target element, including input data actively input by the target element and output data such as audio data and video data corresponding to the target element collected.
[0067] In meeting scenarios, participants often express their needs through significant movements. For example, when participants cannot see the PPT content clearly, they may lean forward or stretch their necks. Conversely, participants with smaller movements are generally considered to not need external assistance during the meeting. This method of using movement as a screening criterion can quickly locate the areas that truly need attention from massive amounts of meeting data, greatly reducing unnecessary data processing.
[0068] Specifically, the system can monitor the movement of target elements within each target recognition area in real time. For example, it can analyze the movement of target elements frame by frame and quantify the target amplitude by calculating parameters such as displacement and angle changes of the target element within a certain time. For example, by recognizing the movement distance of the participant's head or torso, the angle of the arm raised, etc., combined with a preset amplitude threshold, it can determine whether the target amplitude is too large. When the target amplitude of a certain target element exceeds the preset threshold, the target recognition area corresponding to that target element can be marked as a trigger recognition area.
[0069] Once the trigger recognition area is determined, multimodal data for that area can be retrieved. This multimodal data covers multiple dimensions that can be used to infer the target element's needs. Input data comes from content actively input by the target element through the device, such as sending text messages using a conference terminal or issuing commands via a remote control. This data directly reflects the explicitly expressed needs of the target element. Output data mainly consists of audio and video data collected by monitoring devices, such as the participants' speech content, facial expressions, and body movements. Although this data is passively collected, it often contains potential, yet not explicitly expressed, needs of the target element. During data retrieval, the corresponding data segments can be accurately located and extracted based on the range of the trigger recognition area. For video data, the video stream within a specific time period of the trigger recognition area can be obtained. For audio data, the sound signals emitted by the target element within that area can be filtered out. For input data, all information input by the target element through relevant devices within that time period can be retrieved. In this way, the retrieved multimodal data is ensured to be both comprehensive and accurate, providing rich and effective data support for subsequent in-depth analysis of the target element's needs.
[0070] In some embodiments, step S2 includes S21, specifically as follows: S21, identify the target amplitude of the target element in the target identification area, and when the target amplitude is greater than a preset amplitude, use the corresponding target identification area as the trigger identification area.
[0071] Specifically, real-time motion monitoring can be performed on target elements (attendees) within each target recognition area. For example, advanced image recognition and analysis technologies, such as keypoint detection algorithms in computer vision, can be used to analyze the motion of target elements frame by frame. Taking key parts of the attendee's body, such as the head, torso, and limbs, as detection objects, the motion amplitude of the target element can be quantified by calculating parameters such as displacement and angle changes of these parts over a certain period. For instance, by tracking the distance the head moves and the angle at which the arm is raised, combined with a preset calculation formula, the motion is converted into a specific target amplitude value. In a normal meeting scenario, the motion amplitude of attendees is usually within a reasonable range. Within a defined range, for example, when attendees are attentively listening, taking notes, or performing routine relaxation exercises, the movements of key body parts such as the head, torso, and limbs are relatively small and exhibit certain regularity. Through monitoring and data analysis of attendees' movements in numerous real-world meeting scenarios, the range of these natural movements is statistically determined. A preset amplitude threshold is then established. When the amplitude of a identified target element exceeds the preset amplitude, it can be considered that the target element's movement exceeds the range of normal natural activity, potentially indicating that it has encountered a problem during the meeting or requires external assistance. Therefore, the target recognition area corresponding to this target element can be marked as the trigger recognition area. The preset amplitude refers to a threshold pre-set to measure the magnitude of the target element's movement amplitude, based on statistical analysis of the natural movement amplitude of attendees' body parts in normal meeting scenarios.
[0072] The above implementation method can quickly identify areas with potential demand. Compared to comprehensive processing of all data in the entire meeting scenario, this real-time method can reduce unnecessary data processing and improve data processing efficiency.
[0073] S3, perform data parsing on the input data to obtain active adjustment information corresponding to the input data, and identify the dynamic changes of the main elements and accompanying elements in the output data based on the change recognition strategy to obtain the trigger adjustment information of the output data.
[0074] Among them, proactive adjustment information refers to the information obtained after parsing the input data, which directly corresponds to the input data and is used to provide corresponding assistance to the target element. Change recognition strategy refers to the methods and rules used to analyze the dynamic changes of the main element and accompanying elements in the output data. The main element refers to the core object that needs to be focused on in the output data. Its specific meaning varies depending on the type of output data. In audio data, the main element is usually the sound corresponding to the target element (participant), that is, the sound information such as the speech and questions made by the participant. In video data, the main element is generally the target element itself, that is, the image of the participant, including its external manifestations such as its actions, expressions, and postures. Accompanying elements refer to elements that are related to the main element and play an auxiliary role in the analysis of the output data. They also vary depending on the type of output data. In audio data, accompanying elements are background sounds other than the target element's sound, such as ambient noise in the meeting room or the sound of other equipment. These background sounds may affect the clarity and analysis of the main element (the participants' voices). In video data, accompanying elements can be the target element's personal items, such as documents, water cups, and laptops placed on the table by the participants. The state and changes of these items may also reflect the needs of the participants. Trigger adjustment information refers to the information obtained by analyzing the dynamic changes of the main element and accompanying elements in the output data using a change recognition strategy, which is used to trigger adjustments to relevant aspects of the meeting.
[0075] After acquiring multimodal data, the input and output data can be processed separately. For the input data, existing data parsing technology can be used to convert it into active adjustment information corresponding to the input data. For example, if a participant inputs a command to adjust the screen brightness through the device, after parsing, this command can be converted into specific active adjustment information, clearly indicating that the screen brightness needs to be adjusted.
[0076] For output data, a change recognition strategy can be used for analysis. Output data contains main elements and accompanying elements, which may differ in different output data. For example, for audio data, the main element can be the sound corresponding to the target element, and the accompanying element can be other background sounds besides the target element's sound. For video data, the main element can be the target element itself, and the accompanying element can be the target element's personal items. By identifying the dynamic changes of the main and accompanying elements through the change recognition strategy, for example, when participants frequently raise their hands and show confusion, the trigger adjustment information of the output data can be obtained by recognizing and analyzing these dynamic changes. It may be determined that the participants have questions about the meeting content and need staff to answer them. By processing the input and output data, active adjustment information and trigger adjustment information can be obtained respectively. These two types of information can comprehensively reflect the needs of the target element during the meeting.
[0077] In some embodiments, step S3, "identifying the dynamic changes of the main element and accompanying element in the output data based on the change recognition strategy, and obtaining the triggering adjustment information of the output data," includes the following steps: S31, acquire output data, the output data including video data and audio data.
[0078] Specifically, the output data corresponding to the identified trigger recognition area can be obtained, including audio and video data related to that area. This data is continuously collected and stored by the monitoring equipment during the meeting and contains various behavioral information of the target element (participant) within the trigger recognition area. For example, if the trigger recognition area is the area where a participant is located in the meeting room, the video stream corresponding to that area can be extracted from the stored meeting video data, and the sound signal generated in that area can be extracted from the audio recording. By accurately obtaining this data, raw materials are provided for subsequent requirements analysis based on audio and video data, which is the starting point of the entire output data analysis process.
[0079] Among them, video data refers to dynamic image data related to the trigger recognition area that is continuously collected and stored by the monitoring equipment during the meeting, and audio data refers to sound data related to the trigger recognition area that is continuously collected and stored by the monitoring equipment during the meeting.
[0080] S32, based on the dynamic changes of the main elements and accompanying elements in the audio data according to the audio recognition strategy, obtain the trigger adjustment information of the audio data.
[0081] Specifically, audio recognition strategies are analysis methods and rules tailored to the characteristics of audio data. In audio data, the main element is the sound emitted by the target element (participants), such as the content of their speech, tone of voice, and questions. Accompanying elements are other background sounds besides the target element's sound, such as ambient noise in the meeting room and equipment operating sounds. Speech recognition algorithms can analyze the volume, speech rate, tone, and other characteristics of the sound. Different changes in sound characteristics correspond to different demand scenarios. By quantitatively analyzing these characteristics, we can accurately locate the aspects of the meeting that need adjustment and provide participants with more tailored services. For example, a sudden increase in volume may indicate that a participant is coughing, and we may need to provide them with tissues. In addition, by combining changes in background sounds, such as a sudden appearance of noise that may affect the participants' communication, we can determine whether it is necessary to adjust the audio equipment volume or optimize the meeting environment. By comprehensively analyzing the various dynamic changes of the main element and accompanying elements in the audio data, audio recognition strategies can obtain the corresponding trigger adjustment information, such as whether it is necessary to repeat the content to the participants or adjust the meeting audio equipment settings.
[0082] Among them, the audio recognition strategy refers to the strategy of identifying and determining the trigger adjustment information by recognizing audio data.
[0083] Based on the above embodiments, step S32 can be implemented in the following ways: S321, based on timbre recognition, identifies the main sound wave in the audio data as the main element, and the remaining sound waves as accompanying elements.
[0084] Specifically, timbre recognition technology can be used to initially classify audio data. Timbre is the characteristic of sound, and different sound sources (participants, equipment in the meeting environment, and other background sound sources) have unique timbre characteristics. Through a pre-trained timbre recognition model, each sound wave in the audio data is analyzed. This model is trained on a large number of audio samples containing different human voices and environmental sounds, and can accurately identify the sound waves that belong to the participants. These sound waves representing the participants' voices are marked as the main sound waves and identified as the main elements, which contain key information such as participants speaking and asking questions. The remaining sound waves of non-participant voices, such as the sound of the meeting room air conditioner running and the sound of tables and chairs being moved, are classified as accompanying elements. This timbre-based classification lays the foundation for subsequent analysis of the dynamic changes of the main elements and accompanying elements, thereby focusing on the sound information directly related to the needs of the participants, while also taking into account background sounds that may affect the meeting communication. Here, the main sound wave refers to the sound wave corresponding to the participants, the main element refers to the main sound wave in the audio data, and the accompanying element refers to other interfering sound waves in the audio data besides the main sound wave.
[0085] S322, extract the main amplitude of the main element in the audio data, determine that when the main amplitude is greater than the preset main amplitude, take the corresponding sound wave as the main change sound wave, and obtain the main duration of each main change sound wave.
[0086] Specifically, after obtaining the main elements, further analysis can be performed on them. The amplitude information corresponding to the main sound wave can be extracted, i.e., the main amplitude. Amplitude reflects the intensity of sound. The main amplitude is the numerical value of the intensity of the participant's voice. A preset main amplitude can be set in advance. This threshold is determined based on the normal range of sound intensity in meeting communication, combined with actual business needs. When the amplitude of the main sound wave is detected to be greater than the preset main amplitude, it indicates that the participant's voice has changed significantly. These sound waves can be marked as main change sound waves. For example, a participant suddenly raising their voice will trigger this mark. At the same time, the duration of each main change sound wave can be recorded, i.e., the main duration. By obtaining the main change sound waves and their main duration, the specific segments of changes in the intensity of the participant's voice and their duration can be captured, providing key data for subsequent analysis of the participant's emotions and intentions.
[0087] Among them, the main amplitude refers to the amplitude corresponding to the main sound wave, the preset main amplitude refers to the amplitude value set in advance to measure the size of the main amplitude, the main variable sound wave refers to the main sound wave with an amplitude greater than the preset main amplitude, and the main duration refers to the duration corresponding to the main variable sound wave.
[0088] S323, when it is determined that the duration of the main body is less than the preset duration, the corresponding change sound wave of the main body is taken as the instantaneous main body sound wave, and the preset transmission information is retrieved based on the instantaneous main body sound wave.
[0089] After obtaining the main change sound wave and its duration, the duration can be compared with a preset duration. The preset duration is a time threshold set based on common short-term sound change scenarios (such as sudden exclamations or brief questions). If the duration is less than the preset duration, it indicates that the main change sound wave is a short-term sound intensity change, which can be defined as an instantaneous main sound wave. For such instantaneous main sound waves, the corresponding preset transmission information can be retrieved from the preset information database. The preset information database stores common requirement judgments for different instantaneous sound change scenarios. For example, when a short and high-intensity sound is detected, the preset transmission information "Participants may have urgent questions that require a quick response" may be retrieved. This step can quickly react to abnormal sound changes of participants in a short period of time and promptly remind relevant personnel to pay attention to potential needs.
[0090] Among them, the preset duration refers to a time threshold set in advance based on common short-term sound change scenarios (such as a sudden exclamation, a brief question, etc.), the instantaneous main sound wave refers to the main change sound wave with a main duration less than the preset duration, and the preset transmission information refers to the common demand judgment information stored in the preset information database for different instantaneous sound change scenarios.
[0091] S324, when it is determined that the duration of the main body is greater than the preset duration, the corresponding change sound wave of the main body is taken as the continuous main body sound wave, the text information corresponding to the continuous main body sound wave is extracted, and annotation adjustment information is generated based on the text information.
[0092] Specifically, when the duration of the main body exceeds the preset duration, it indicates that the main body change sound wave is a long-lasting change in sound intensity, which can be identified as a continuous main body sound wave. This long-lasting, high-intensity sound wave may be considered to be an important part of the meeting. Speech recognition technology can be used to convert the continuous main body sound wave into text information. For example, the content of continuous speeches by participants can be converted into readable text. Then, semantic analysis can be performed on this text information to understand the content and intentions expressed by the participants. Based on the analysis results, corresponding annotation and adjustment information can be generated to highlight the relevant meeting content, making it easier for other participants to view the key meeting content later.
[0093] Among them, the continuous main sound wave refers to the main sound wave with a duration longer than the preset duration, the text information refers to the readable text content converted from the continuous main sound wave using speech recognition technology, and the annotation adjustment information refers to the specific information generated by semantic analysis of the text information after the continuous main sound wave is converted, to understand the content and intention expressed by the participants, and to indicate the key points of the corresponding meeting content. The annotation adjustment information can guide relevant personnel (such as meeting recorders, participants, etc.) to quickly locate and view the key content in the meeting, and facilitate the organization, review and utilization of key meeting information.
[0094] S325, extract the accompanying amplitude of the accompanying element in the audio data, determine that when the accompanying amplitude is greater than the preset accompanying amplitude, take the corresponding sound wave as the accompanying change sound wave, and obtain the accompanying duration of each accompanying change sound wave.
[0095] Specifically, after analyzing the main elements, the accompanying elements can be analyzed accordingly. The amplitude of the accompanying sound waves can be extracted, which represents the intensity of the background sound in the meeting environment. Similarly, a preset accompanying amplitude can be set to measure whether the background sound has abnormally increased. When the accompanying amplitude is greater than the preset accompanying amplitude, it means that the intensity of the background sound in the meeting environment exceeds the normal range. These sound waves can be marked as accompanying change sound waves and their duration can be recorded. For example, if a loud equipment failure sound suddenly occurs in the meeting room, the corresponding equipment failure sound can be marked as an accompanying change sound wave and its duration recorded. This step can promptly detect sound changes in the meeting environment that may interfere with normal communication and provide data support for subsequent countermeasures.
[0096] Among them, the accompanying amplitude refers to the amplitude of the accompanying sound wave, the preset accompanying amplitude is a set value used to measure whether the background sound has abnormally increased, the accompanying change sound wave is the sound wave that is marked when the accompanying amplitude is greater than the preset accompanying amplitude, and the accompanying duration refers to the duration of the accompanying change sound wave.
[0097] S326, when it is determined that the accompanying duration is less than the preset duration, the corresponding accompanying change sound wave is taken as the instantaneous accompanying sound wave, and the processing and adjustment information is retrieved based on the instantaneous accompanying sound wave.
[0098] Specifically, after obtaining the accompanying duration, it can be compared with a pre-set preset duration. If the accompanying duration is shorter than the preset duration, it indicates that the accompanying sound wave is a short-term background sound anomaly, defined as an instantaneous accompanying sound wave. For instantaneous accompanying sound waves, corresponding handling and adjustment information can be retrieved from a preset handling plan library. For example, if the brief noise is caused by a temporary equipment malfunction, the handling and adjustment information of "checking relevant equipment to confirm whether repair is needed" can be retrieved. By quickly responding to instantaneous background sound changes, sudden environmental problems that may affect the meeting can be dealt with in a timely manner, maintaining the normal order of the meeting. Here, an instantaneous accompanying sound wave refers to an accompanying sound wave with an accompanying duration shorter than the preset duration, and the handling and adjustment information refers to the specific information formulated for instantaneous accompanying sound waves to guide the corresponding handling measures.
[0099] S327, when it is determined that the accompanying duration is greater than the preset duration, the corresponding accompanying change sound wave is taken as the continuous accompanying sound wave, and an elimination sound wave opposite to the continuous accompanying sound wave is generated, and the elimination sound wave is used as elimination adjustment information.
[0100] Specifically, when the duration of the accompanying noise exceeds the preset duration, the background noise can be considered to be in an abnormal state for an extended period and can be identified as a continuous accompanying sound wave. In order to eliminate the impact of prolonged abnormal background noise on meeting communication, audio processing technology can be used to generate a cancellation sound wave that is opposite to the continuous accompanying sound wave in terms of frequency, amplitude, and other characteristics. For example, for continuous low-frequency noise, a low-frequency sound wave with the opposite phase can be generated to cancel it out. The generated cancellation sound wave is sent as cancellation adjustment information to the relevant audio equipment, which can activate the corresponding noise reduction program, reduce the interference of background noise on the meeting, and create a clearer communication environment for the participants.
[0101] Among them, the continuous accompanying sound wave refers to the accompanying sound wave with a duration longer than the preset duration, that is, the background sound is in an abnormal state for a long time. The elimination sound wave refers to the use of audio processing technology to generate a sound wave with the opposite characteristics in frequency, amplitude and other aspects to the continuous accompanying sound wave, in order to cancel the continuous accompanying sound wave and reduce its influence. The elimination adjustment information refers to the generated elimination sound wave.
[0102] S33, based on the dynamic changes of the main elements and accompanying elements in the video data according to the video recognition strategy, obtain the trigger adjustment information of the video data.
[0103] Specifically, video recognition strategies are analytical strategies designed for video data. In video data, the main element is the target element (the attendee) themselves, including their actions, expressions, postures, and other external manifestations. Accompanying elements are the attendee's personal belongings, such as the placement of documents on the table and water glasses. Computer vision technology can be used to analyze video data frame by frame. For example, human posture recognition algorithms can detect attendees' body movements; frequent hand raising might indicate an intention to speak, while leaning forward and frowning might suggest difficulty seeing the screen content. Expression recognition algorithms can analyze attendees' facial expressions; for example, a confused expression might indicate... If participants don't understand the meeting content, they can be reminded to have the relevant personnel explain it. Simultaneously, changes in accompanying elements can be observed. For example, if participants repeatedly look at their water glasses, it might mean they need to rehydrate; if documents are frequently flipped through, they might be searching for relevant information, in which case a clearer presentation of the meeting materials might be needed. By comprehensively considering the dynamic changes of both main and accompanying elements in the video data, and using video recognition strategies for reasoning and judgment, triggering adjustment information can be obtained from the video data. This includes determining whether the projection screen needs adjustment or whether beverages should be provided to participants, thereby achieving accurate capture and response to the participants' potential needs.
[0104] Among them, video recognition strategy refers to the strategy of determining the trigger adjustment information corresponding to the target element by analyzing video data.
[0105] Based on the above embodiments, step S33 can be implemented in the following ways: S331, identify the target element in the video data as the main element, and the accompanying placement element as the accompanying element.
[0106] Specifically, video data can be analyzed, and object detection algorithms can be used to accurately locate attendees (target elements) in the video frame and identify them as main elements. These algorithms are trained on a large number of image samples containing human features, enabling them to quickly and accurately identify attendees in the frame and outline their contours. At the same time, they can identify accompanying elements within the target recognition area corresponding to the target element in the video frame, such as documents, water cups, and laptops located within the target recognition area, and identify these items as accompanying elements. During the recognition process, the algorithm can analyze the shape, color, texture, and other features of the items and compare them with a pre-stored item feature database, thereby achieving accurate identification of various items. By clearly distinguishing between main elements and accompanying elements, a foundation is laid for subsequent analysis of the dynamic changes of both, thus enabling the targeted capture of visual information related to the needs of attendees.
[0107] Here, the main element refers to the target element in the video data, and the accompanying element refers to the accompanying placement element in the video data.
[0108] S332, Identify the change state of the corresponding element part of the main element in the video data, and generate instruction adjustment information based on the change state.
[0109] Specifically, the system can monitor the changes in the movements and postures of various body parts (elements) of attendees in real time using video data. Human posture recognition algorithms, such as OpenPose, can detect multiple key joints in the human body. By analyzing the positional changes and relative relationships of these joints, the system can determine the attendees' body movements, such as raising their hands, bending over, or leaning forward. Simultaneously, combined with facial expression recognition algorithms, the system analyzes changes in the attendees' facial expressions, identifying expressions such as smiling, frowning, and confusion. For example, if an attendee is detected leaning forward and focusing their gaze on the projection screen for an extended period, it can be inferred that they may not be able to see the screen content clearly. This allows for the generation of instructions to "adjust the projection image clarity or enlarge the font." Through precise identification and analysis of the changes in the various parts of the main elements, the system can translate the attendees' body language and facial expressions into specific service instructions, enabling timely responses to the attendees' explicit needs.
[0110] Among them, the element part refers to the various parts of the main element (such as the participants) in the video data, the change state refers to the changes in the movements, postures, etc. of the various element parts of the participants' bodies, and the instruction adjustment information refers to the specific information generated based on the identification and analysis results of the change state of the element parts of the participants, which is used to instruct for corresponding adjustments and operations.
[0111] Based on the above embodiments, step S332 can be implemented in the following ways: S3321, Identify the change state of the corresponding element part of the main element in the video data, the change state includes independent change and combined change.
[0112] Specifically, computer vision algorithms can be used to analyze video data frame by frame, monitoring changes in various parts of the participants' bodies (elemental parts) in real time. Through human posture recognition algorithms (such as OpenPose) and facial expression recognition algorithms, the movement, posture, and facial expression changes of each elemental part can be accurately captured. These changes can be classified. Independent changes refer to changes occurring in a single elemental part, such as a single action like raising an arm or turning the head. Combined changes are changes occurring in conjunction with multiple elemental parts, such as covering the mouth with the hand when yawning, or a combination of actions produced by multiple parts of the body leaning forward while frowning and looking at the screen. By analyzing and judging the changes in elemental parts, these two types of changes, namely independent changes and combined changes, can be clearly distinguished, laying the foundation for subsequent targeted analysis strategies.
[0113] Among them, the change state refers to the changes that occur in the corresponding parts of the main element in terms of action, posture and expression in the video data. Independent change refers to the change that occurs in a single part of the main element, which involves only a single body part making a change in action, posture or expression. Combined change refers to the changes that occur in multiple parts of the main element in a coordinated manner, which is a combination of actions, postures or expressions produced by the cooperation of multiple body parts.
[0114] S3322, when the change state is determined to be an independent change, the change range of the element part is identified, and when the change range of the part is determined to be greater than the preset part range, the corresponding preset instruction information is retrieved based on the element part as instruction adjustment information.
[0115] When the change is determined to be an independent change, the magnitude of the change in the element can be further quantified. By calculating parameters such as displacement and angle changes of the element over a certain period of time, the magnitude of the change can be obtained. Then, a preset magnitude can be set in advance. This threshold is based on statistical analysis of the natural range of motion of the body parts of participants in a normal meeting scenario. When the magnitude of the change in the element is detected to be greater than the preset magnitude, it indicates that the action may have a specific intention. At this time, the corresponding preset instruction information can be retrieved from the preset instruction information library based on the element. For example, if it is detected that the participant's head is tilted back at a large angle and maintained for a long time, it can be determined that the participant may be feeling tired. The preset instruction information "provide the participant with a cup of coffee" can be retrieved from the instruction information library and used as the instruction adjustment information to achieve timely response to the needs of the participants.
[0116] Among them, the change amplitude of the part refers to the degree of change in the action and posture of the main element (such as the participants) in the video data when the part of the element changes independently; the preset part amplitude refers to the amplitude threshold set in advance based on the statistical analysis of the natural movement amplitude of the body parts of the participants in a normal meeting scenario; the preset instruction information refers to the corresponding operation instruction information set in advance in the instruction information library for different element parts and their possible changes; and the instruction adjustment information refers to the corresponding preset instruction information retrieved from the preset instruction information library according to the element part when the change amplitude of the detected part part is greater than the preset part amplitude.
[0117] S3323, when the change state is determined to be a combination change, the corresponding element part is taken as the combination part, and when the combination parts have an intersection, the preset instruction information corresponding to the corresponding combination part is taken as the instruction adjustment information.
[0118] Specifically, when a change is determined to be a combined change, multiple elements that change synergistically can be considered as combined parts. Then, by analyzing the spatial relationships and synergistic actions among these combined parts, it can be determined whether the combined parts have any overlap. For example, when a participant squints while covering their eyes with their hand, the hand and eyes can be considered as combined parts. Their changes are interconnected in time and space, jointly indicating that the light may be a bit glaring at this moment. At this time, based on the characteristics of the combined parts, the corresponding preset instruction information, such as "Please close the curtains," can be retrieved from the preset instruction information database and used as the instruction adjustment information. By analyzing the combined changes, the complex demand signals of the participants can be understood more comprehensively and accurately, thereby providing service instructions that are more in line with actual needs.
[0119] Among them, the combined part refers to the multiple element parts that change in action, posture or expression when the change state of the corresponding element part of the main element (such as the participants) in the video data is judged as a combined change.
[0120] S333, Obtain the display attributes of accompanying elements in the video data, wherein the display attributes include transparency attributes and concealment attributes.
[0121] Specifically, the display attributes of accompanying elements can be determined by analyzing their visual presentation in the video frame. The transparency attribute corresponds to items with transparent materials, such as a glass water cup, whose outline and internal state (such as the water level) can be directly identified through the video. The concealment attribute corresponds to items with opaque materials, such as a stainless steel thermos. Here, display attributes refer to the visual characteristics of accompanying elements in the video data. They can be used to describe the observable and identifiable state of accompanying elements in the video frame. The transparency attribute refers to the display attribute of accompanying elements whose materials are transparent, and whose outline and internal state can be clearly identified and observed directly through the video. The concealment attribute refers to the display attribute of accompanying elements whose materials are opaque, and whose outline and internal state cannot be clearly identified and observed directly through the video.
[0122] S334, when the display attribute is determined to be transparent, the corresponding accompanying element is taken as an identifiable element, the usage height of the identifiable element is identified, and when the usage height is less than the limit height, displacement adjustment information is generated.
[0123] Specifically, when the display attribute of an accompanying element is transparent, it can be identified as a recognizable element. Further quantitative analysis can determine the usage height corresponding to the recognizable element. For example, taking a water cup as an example, by measuring the vertical position of the remaining water in the cup in the video frame, combined with the known frame ratio and actual scene size, its usage height (i.e., the height of the remaining water in the actual scene) can be calculated. Then, a limit height can be preset. When it is determined that the usage height of the recognizable element is less than the limit height, for example, when the height corresponding to the remaining water is significantly lower, it is speculated that the water in the cup is about to run out. At this time, replacement adjustment information can be generated, such as "replenish the beverage for this participant". By monitoring and analyzing the usage height of recognizable accompanying elements, the potential item usage needs of participants can be captured, and corresponding services can be provided in a timely manner.
[0124] Among them, the identifiable element refers to the accompanying element with the transparent attribute; the usage height refers to the vertical height of the internal material (such as water in a cup) of the identifiable element in the actual scene; the limit height refers to a pre-set threshold value used to measure whether the usage height of the identifiable element has reached the point where action needs to be taken; and the placement adjustment information refers to the instruction information generated according to the specific situation when the usage height of the identifiable element is less than the limit height, which is used to indicate the corresponding item replacement or replenishment operations.
[0125] S335, when it is determined that the display attribute is a hidden attribute, the corresponding accompanying element is treated as an unrecognizable element, the usage tilt of the unrecognizable element is identified, and when it is determined that the usage tilt is greater than the limit tilt, adjustment information is generated.
[0126] Specifically, if the display attribute of an accompanying element is hidden, it can be marked as an unrecognizable element. The outline and posture of the unrecognizable element can be identified through edge detection, shape fitting and other technologies, and then its tilt angle (i.e. the angle between the item and the horizontal or vertical direction) can be calculated. Similarly, a limit tilt angle can be preset to measure whether the tilt state of the item is abnormal. When it is determined that the tilt angle of the unrecognizable element is greater than the limit tilt angle, such as when the tilt angle of an opaque water cup is greater than the limit tilt angle, it can be inferred that the corresponding water cup may be empty. Corresponding adjustment information can be generated to remind staff to add water to the corresponding participants.
[0127] Among them, unidentifiable elements refer to accompanying elements whose display attribute is hidden; tilt angle refers to the angle that can be used to reflect the tilt state of the unidentifiable element; limit tilt angle refers to a pre-set critical value used to measure whether the tilt angle of the unidentifiable element has reached an abnormal level; and add adjustment information refers to the instruction information generated based on the predicted possible situation when the tilt angle of the unidentifiable element is greater than the limit tilt angle, which is used to instruct the corresponding item addition operation.
[0128] The above implementation method can comprehensively capture the demand signals of participants, including the demands explicitly expressed through input data and the potential demands expressed through output data, thereby improving the accuracy and comprehensiveness of participant demand identification.
[0129] S4, send the active adjustment information and the triggered adjustment information to the adjustment device of the target element.
[0130] After receiving both proactive and triggered adjustment information, the corresponding adjustment information can be sent to the adjustment device corresponding to the target element. Each target element has a corresponding adjustment device, which can be monitored by the venue's management personnel. When the target element's adjustment device receives the corresponding adjustment information, the management personnel can process it accordingly and provide assistance to meet the needs of the target element. For example, if the adjustment device receives the adjustment information "replace the drinking water for this person," the management personnel can provide new drinking water to the target element. Here, "adjustment device" refers to equipment capable of processing the analyzed adjustment information.
[0131] See Figure 3 This is a schematic diagram of the structure of a mobile edge intelligent inference system supporting multimodal input provided by an embodiment of the present invention. The data processing system based on the mobile edge intelligent inference system supporting multimodal input includes: The stitching module is used to respond to the pre-completion information, receive the target image collected by the monitoring device on the target area, identify the static planning area and dynamic stitching area of each target element in the target image and stitch them together to obtain the target recognition area of the target element; The retrieval module is used to retrieve multimodal data of the trigger recognition area based on the target amplitude of the target element in the target recognition area, and the corresponding target recognition area is used as the trigger recognition area. The multimodal data includes input data and output data. The parsing module is used to parse the input data to obtain active adjustment information corresponding to the input data, and to identify the dynamic changes of the main elements and accompanying elements in the output data based on the change recognition strategy to obtain the trigger adjustment information of the output data. The sending module is used to send the active adjustment information and the triggered adjustment information to the adjustment device of the target element.
[0132] Figure 4 The apparatus of the illustrated embodiment can be used to perform corresponding actions. Figure 1 The steps in the method embodiments shown are implemented in a similar manner and have similar technical effects, and will not be repeated here.
[0133] See Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. The electronic device 40 includes: a processor 41, a memory 42, and a computer program; wherein... The memory 42 is used to store the computer program, and the memory may also be flash memory. The computer program is, for example, an application program or functional module that implements the above method.
[0134] The processor 41 is configured to execute the computer program stored in the memory to implement the various steps performed by the device in the above method. For details, please refer to the relevant descriptions in the preceding method embodiments.
[0135] Alternatively, the memory 42 can be either standalone or integrated with the processor 41.
[0136] When the memory 42 is a device independent of the processor 41, the device may further include: Bus 43 is used to connect the memory 42 and the processor 41.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A mobile edge intelligent inference method supporting multimodal input, characterized in that, include: In response to the pre-completion information, the system receives the target image collected by the monitoring equipment for the target area, identifies and stitches together the static planning area and dynamic stitching area of each target element in the target image to obtain the target recognition area of the target element. Based on the target amplitude of the target element in the target recognition area, the corresponding target recognition area is used as the trigger recognition area, and the multimodal data of the trigger recognition area is retrieved. The multimodal data includes input data and output data. The input data is parsed to obtain active adjustment information corresponding to the input data. Based on the change recognition strategy, the dynamic changes of the main element and the accompanying element in the output data are identified to obtain the trigger adjustment information of the output data. The active adjustment information and the triggered adjustment information are sent to the adjustment device of the target element; The process of identifying and stitching together the static planning region and dynamic stitching region of each target element in the target image to obtain the target recognition region of the target element includes: Identify the positional contours of static location elements and the placement contours of static placement elements in the target image. The target image refers to the image obtained by capturing images of the target area where the meeting is being held. The static location elements refer to the chairs where the participants are sitting, and the static placement elements refer to the conference table. Determine the position contour that is closest to each placement side length in the placement contour as the processing contour for each placement side length; The fixed side length and extended side length in the processing contour are processed to obtain the static planning area of the target element; Based on the placement side length, the corresponding static planning area is mirrored to obtain the copied planning area; The copy planning area is dynamically filled in according to the accompanying outline of the accompanying placement element in the placement outline to obtain the dynamic splicing area of the target element. The static planning area and the dynamic splicing area are spliced together to obtain the target recognition area of the target element.
2. The method according to claim 1, characterized in that, The process of processing the fixed side length and extended side length in the processing contour to obtain the static planning area of the target element includes: The side length parallel to the corresponding placement side length in the processing contour is taken as the fixed side length, and the remaining side length is taken as the extended side length. Delete the fixed side length closest to the corresponding placement side length, and extend the extended side length to one side of the corresponding placement side length until it intersects with the placement side length, thus obtaining the static planning area of the target element.
3. The method according to claim 1, characterized in that, The step of dynamically filling the copy planning area according to the accompanying outline of the accompanying placement element in the placement outline to obtain the dynamic splicing area of the target element includes: Identify the accompanying contours of the accompanying placement elements in the placement contour, take the accompanying contours that intersect with the copy planning area as the merged contours, and take the remaining accompanying contours as the contours to be selected. The first combined contour is obtained by the union of the merged contour and the corresponding replicated planning area; The target image is processed by coordinate conversion to obtain the straight-line distance between the center point of the contour to be selected and the first combined contour. The contour to be selected whose straight-line distance is less than the preset filtering distance is selected as the corresponding first combined contour. Obtain the extreme coordinate values of the selected contour and the first combined contour, construct a reconstructed rectangle based on the extreme coordinate values, and use the reconstructed rectangle as the dynamic splicing area of the target element.
4. The method according to claim 1, characterized in that, The step of using the target recognition area as the trigger recognition area based on the target amplitude of the target element in the target recognition area includes: Identify the target amplitude of the target element in the target identification area, and when the target amplitude is greater than a preset amplitude, use the corresponding target identification area as the trigger identification area.
5. The method according to claim 1, characterized in that, The change recognition strategy identifies the dynamic changes of the main elements and accompanying elements in the output data to obtain the triggering adjustment information of the output data, including: Acquire output data, which includes video data and audio data; Based on the dynamic changes of the main elements and accompanying elements in the audio data using the audio recognition strategy, the triggering and adjustment information of the audio data is obtained. Based on the video recognition strategy, the dynamic changes of the main elements and accompanying elements in the video data are analyzed to obtain the trigger adjustment information of the video data.
6. The method according to claim 5, characterized in that, The method based on audio recognition strategy to dynamically change the main elements and accompanying elements in audio data, and to obtain the trigger adjustment information of audio data, includes: Based on timbre recognition, the main sound wave in the audio data is taken as the main element, and the remaining sound waves are taken as accompanying elements; Extract the main amplitude of the main element in the audio data, and when the main amplitude is greater than the preset main amplitude, take the corresponding sound wave as the main change sound wave, and obtain the main duration of each main change sound wave; When the duration of the main body is determined to be less than the preset duration, the corresponding change in the main body sound wave is taken as the instantaneous main body sound wave, and the preset transmission information is retrieved based on the instantaneous main body sound wave; When the duration of the main body is determined to be greater than the preset duration, the corresponding change sound wave of the main body is taken as the continuous main body sound wave, the text information corresponding to the continuous main body sound wave is extracted, and annotation adjustment information is generated based on the text information; Extract the accompanying amplitude of the accompanying element in the audio data. When the accompanying amplitude is greater than the preset accompanying amplitude, the corresponding sound wave is taken as the accompanying change sound wave, and the accompanying duration of each accompanying change sound wave is obtained. When it is determined that the accompanying duration is less than the preset duration, the corresponding accompanying change sound wave is taken as the instantaneous accompanying sound wave, and the processing and adjustment information is retrieved based on the instantaneous accompanying sound wave. When the duration of the accompanying sound wave is determined to be longer than a preset duration, the corresponding accompanying sound wave is taken as a continuous accompanying sound wave, and a cancellation sound wave opposite to the continuous accompanying sound wave is generated. The cancellation sound wave is used as cancellation adjustment information.
7. The method according to claim 5, characterized in that, The step of obtaining trigger adjustment information for video data based on the dynamic changes of main elements and accompanying elements in the video data according to the video recognition strategy includes: Identify target elements in video data as main elements, and accompanying elements as accompanying elements; Identify the changing state of the corresponding element parts of the main element in the video data, and generate instruction adjustment information based on the changing state; Obtain the display attributes of accompanying elements in the video data, including transparency and concealment attributes; When the display attribute is determined to be transparent, the corresponding accompanying element is taken as an identifiable element, the usage height of the identifiable element is identified, and when the usage height is less than the limit height, displacement adjustment information is generated. When the display attribute is determined to be a hidden attribute, the corresponding accompanying element is treated as an unrecognizable element. The usage tilt of the unrecognizable element is identified. If the usage tilt is greater than the limit tilt, adjustment information is generated.
8. The method according to claim 7, characterized in that, The method of identifying the changing state of the corresponding element parts of the main element in the video data, and generating instruction adjustment information based on the changing state, includes: Identify the changing state of the corresponding element parts of the main element in the video data, wherein the changing state includes independent changes and combined changes; When the change state is determined to be an independent change, the change range of the element part is identified. When the change range of the part is determined to be greater than the preset part range, the corresponding preset instruction information is retrieved based on the element part as instruction adjustment information. When the change state is determined to be a combination change, the corresponding element part is taken as the combination part. When the combination parts have an intersection, the preset instruction information corresponding to the corresponding combination part is taken as the instruction adjustment information.
9. A mobile edge intelligent inference system supporting multimodal input according to any one of claims 1-8, characterized in that, include: The stitching module is used to respond to the pre-completion information, receive the target image collected by the monitoring device on the target area, identify the static planning area and dynamic stitching area of each target element in the target image and stitch them together to obtain the target recognition area of the target element; The retrieval module is used to retrieve multimodal data of the trigger recognition area based on the target amplitude of the target element in the target recognition area, and the corresponding target recognition area is used as the trigger recognition area. The multimodal data includes input data and output data. The parsing module is used to parse the input data to obtain active adjustment information corresponding to the input data, and to identify the dynamic changes of the main elements and accompanying elements in the output data based on the change recognition strategy to obtain the trigger adjustment information of the output data. The sending module is used to send the active adjustment information and the triggered adjustment information to the adjustment device of the target element.
Citation Information
Patent Citations
System and method for processing visual, auditory, olfactory, and / or haptic information
CN103003761A