Language Model Dynamic Processing Method, System and Medium

By dividing the image content and generating path recording layer, and determining the dynamic interactive content based on the user's dynamic interaction behavior, the problem that the existing technology cannot dynamically capture user interaction behavior and generate flexible text is solved, and efficient and accurate information acquisition and interactive experience are achieved.

CN119850897BActive Publication Date: 2025-06-17ZHEJIANG ELECTRIC POWER TRADING CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510335113.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-17
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Existing visual processing technologies are difficult to dynamically capture the user's interaction behavior on images, and the generated text lacks the flexibility to match the user's personalized operations, and cannot meet the user's needs for dynamic interaction.

Method used

By performing content slicing processing on the image, multiple content areas are generated, path recording layers corresponding to the image are generated, dynamic interactive content is determined based on the user's dynamic interaction behavior, and corresponding language results are output.

Benefits of technology

It realizes intelligent generation of dynamic data in combination with user dynamic interaction behavior, improves the accuracy and efficiency of information acquisition, and enhances the interactive experience between users and images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850897B_ABST
    Figure CN119850897B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system and medium for dynamically processing a language model, relating to data processing technology. The method includes: after determining that a user selects to perform dynamic processing of a language model on a first file in a preset modality, performing content segmentation processing on the first file to generate a plurality of content regions, where the first file is an image; generating a path record layer corresponding to the first file; after determining that the user triggers any pixel point in the path record layer, determining dynamic interaction content based on the spatio-temporal relationship of the triggered pixel point; determining at least one content or at least one content region based on the dynamic interaction content and outputting a language result. The present invention can intelligently generate dynamic data in combination with the dynamic interaction behavior of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data processing card technology, and in particular, to a method, system, and medium for dynamically processing a language model. Background Art

[0002] Driven by modern information technology, visual information processing has become the focus of attention in various industries. With the popularization of intelligent devices, users expect to extract the required information from complex image data through simple interaction methods. For example, in the field of autonomous driving, the system needs to analyze road images in real time to make decisions; in e-commerce, users hope to obtain detailed descriptions of products through pictures. This trend has prompted developers to explore more intelligent image understanding and natural language generation technologies to enhance the interactive experience between images and users.

[0003] Existing visual processing technologies mainly rely on static image recognition and predefined text generation. Although these technologies have made progress in some aspects, there are still many deficiencies. For example, traditional methods are usually limited to recognizing certain specific objects in images and are difficult to dynamically capture the interactive behavior of users on images. In addition, the generated text is often in a fixed pattern and lacks the flexibility to match the user's personalized operations. This limitation cannot meet the user's demand for dynamic interaction.

[0004] Therefore, how to combine the dynamic interaction behavior of users to intelligently generate dynamic data has become an urgent problem to be solved. Summary of the Invention

[0005] Embodiments of the present invention provide a method, system, and medium for dynamically processing a language model, which can combine the dynamic interaction behavior of users to intelligently generate dynamic data.

[0006] In the first aspect of the embodiments of the present invention, a method for dynamically processing a language model is provided, including:

[0007] After determining that the user selects to perform dynamic processing of the language model on a first file of a preset modality, the first file is segmented according to its content to generate a plurality of content regions, and the first file is an image;

[0008] Generate a path record layer corresponding to the first file;

[0009] After determining that the user triggers any pixel point in the path record layer, determine the dynamic interaction content based on the spatio-temporal relationship of the triggered pixel point;

[0010] Determine at least one content or at least one content region based on the dynamic interaction content and output a language result.

[0011] Optionally, the segmenting the first file according to its content to generate a plurality of content regions includes:

[0012] Perform graphic comparison on the first file based on a preset element database to determine the first image elements that pass the comparison within the first file;

[0013] After coordinate processing the first file, obtain the coordinate intervals of each first image element, and generate the content coordinate intervals of the first image elements after processing the coordinate intervals;

[0014] Take all pixel points within the content coordinate intervals as the corresponding content areas.

[0015] Optionally, the generating the content coordinate intervals of the first image elements after processing the coordinate intervals includes:

[0016] Obtain the extreme values of the abscissa and the extreme values of the ordinate of the coordinate intervals;

[0017] Increase the maximum value among the extreme values of the abscissa and the maximum value among the extreme values of the ordinate by a preset value respectively, and decrease the minimum value among the extreme values of the ordinate and the minimum value among the extreme values of the ordinate by a preset value respectively to obtain the content coordinate intervals.

[0018] Optionally, the generating the path recording layer corresponding to the first file includes:

[0019] Obtain the coordinates of each pixel point in the first file;

[0020] Establish layer coordinates corresponding to the coordinates of each pixel point to generate a path recording layer.

[0021] Optionally, the determining the dynamic interaction content based on the spatio-temporal relationship of the triggered pixel point after determining that the user triggers any pixel point in the path recording layer includes:

[0022] Real-time obtain the first pixel point triggering the path recording layer, and extract all the second pixel points triggered historically within a preset time period;

[0023] Generate a pixel point sequence based on the time sequence of the first pixel point and the second pixel points, and sort the first pixel point and the second pixel points according to the time from near to far based on the pixel point sequence, so that the first pixel point and the second pixel points respectively have corresponding sorting numbers;

[0024] Perform tracing processing on the first pixel point and the second pixel points in order based on the sorting numbers to determine the dynamic interaction content.

[0025] Optionally, the performing tracing processing on the first pixel point and the second pixel points in order based on the sorting numbers to determine the dynamic interaction content includes:

[0026] Connect the first pixel point and the second pixel points in order based on the sorting numbers to generate a tracing route;

[0027] Extract the figure formed by the traced route and compare it with the preset instruction database to determine the first instruction that passes the comparison. Each instruction has a preset response strategy;

[0028] Extract the dynamic interaction content corresponding to this instruction based on the first instruction.

[0029] Optionally, the extracting the dynamic interaction content corresponding to this instruction based on the first instruction includes:

[0030] If the preset response strategy of the first instruction is a positioning strategy, then extract the abscissa extreme values and ordinate extreme values of all the first pixel points and second pixel points in the traced route;

[0031] Determine the abscissa intermediate value and ordinate intermediate value based on the abscissa extreme values and ordinate extreme values to obtain a positioning intermediate point;

[0032] Determine the initial segmentation area of the path recording layer where the positioning intermediate point is located, and output the corresponding output language result according to the initial segmentation area where the positioning intermediate point is located.

[0033] Optionally, the initial segmentation area of the path recording layer is obtained through the following steps, including:

[0034] Obtain the number of horizontal pixel points and the number of vertical pixel points in the path recording layer;

[0035] Equally divide the number of horizontal pixel points according to the preset horizontal equal division value to obtain multiple horizontal equal division points, and equally divide the number of vertical pixel points according to the preset vertical equal division value to obtain multiple vertical equal division points;

[0036] Generate equidistant lines based on the horizontal equal division points and vertical equal division points to equally divide the path recording layer, obtain the initial segmentation area of the path recording layer, and add corresponding language orientation information to each initial segmentation area. The output language result includes the language orientation information.

[0037] Optionally, the extracting the dynamic interaction content corresponding to this instruction based on the first instruction includes:

[0038] If the preset response strategy of the first instruction is a content recognition strategy, then extract the abscissa extreme values and ordinate extreme values of all the first pixel points and second pixel points in the traced route;

[0039] Generate a selected area based on the abscissa extreme values and ordinate extreme values;

[0040] Map the pixel points within the selected area to the first file to determine the dynamic interaction content.

[0041] Optionally, determining at least one content or at least one content area output language result based on the dynamic interaction content includes:

[0042] Obtain a first image element whose pixel point coordinate position intersects with the selected area, determine the number of pixel points in the intersection to obtain a first intersection number, and count the total number of pixel points in the first image element;

[0043] Calculate the ratio of the first intersection number to the first total number to obtain a first proportion. If the first proportion is greater than a preset value, use the content of the corresponding first image element as the output language result;

[0044] If it is determined that all first proportions are less than or equal to the preset value, determine a content area to output the language result.

[0045] Optionally, determining a content area to output the language result includes:

[0046] Obtain a content area whose pixel point coordinate position intersects with the selected area, and determine the number of pixel points in the intersection to obtain a second intersection number;

[0047] If it is determined that there is only one content area with an intersection with the selected area, determine the corresponding content area to output the language result;

[0048] If it is determined that there are multiple content areas with intersections with the selected area, select the content area with the larger second intersection number to output the language result, and the output language result includes all the content in the content area.

[0049] In the second aspect of the embodiments of the present invention, a language model dynamic processing system is provided, including:

[0050] A judgment module, configured to, after judging that the user selects to perform language model dynamic processing on a first file in a preset modality, perform content segmentation processing on the first file to generate multiple content areas, where the first file is an image;

[0051] A generation module, configured to generate a path record layer corresponding to the first file;

[0052] A determination module, configured to, after judging that any pixel point in the path record layer is triggered by the user, determine dynamic interaction content based on the spatio-temporal relationship of the triggered pixel point;

[0053] An output module, configured to determine at least one content or at least one content area output language result based on the dynamic interaction content.

[0054] In a third aspect of the embodiments of the present invention, a storage medium is provided, in which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the method described in the first aspect and various possible designs of the first aspect of the present invention.

[0055] By performing graphic comparison on the first file based on a preset element database, the present invention can accurately determine the first image element within the image, then generate a content coordinate interval through coordinate processing, and further determine the pixel points within the content coordinate interval as the content area. This series of operations enables the image to be finely divided into multiple content areas, and users can interact with the path recording layer to obtain information about the area of interest in a targeted manner. It improves the accuracy and efficiency of information acquisition, and effectively solves the problem that the traditional technology cannot accurately locate the focus of user attention.

[0056] The present invention obtains the first pixel point in the trigger path recording layer in real time, extracts the second pixel points triggered historically within a preset time period, generates a pixel point sequence according to the time sequence and sorts it, and then determines the first instruction through trace processing and comparison with a preset instruction database. This process can deeply understand the user's operation intention and generate corresponding language results according to different instruction response strategies. For example, when the user draws a circular trajectory on the image, the system can accurately recognize it as an "content recognition" instruction and provide accurate language feedback accordingly, greatly enhancing the interaction experience between the user and the image and making up for the deficiency of the existing system in understanding the user's interaction behavior.

[0057] When determining the output language result, for the content recognition strategy, the system obtains the first image element whose pixel point coordinates intersect with the selected area, calculates the ratio of the number of intersection pixel points to the total number of pixel points within the first image element, and uses this to judge whether to use the content of the corresponding image element as the output language result; if all ratios do not meet the conditions, a content area is determined to output the language result. This way of analyzing the intersection of the image element and the selected area and comprehensively judging the content area can accurately output the language result closely related to the user's operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a schematic flowchart of a method for dynamically processing a language model provided by an embodiment of the present invention;

[0059] Figure 2 is a schematic diagram of a trace route provided by an embodiment of the present invention;

[0060] Figure 3 is a schematic structural diagram of a system for dynamically processing a language model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] See Figure 1 , which is a schematic flowchart of a method for dynamically processing a language model provided by an embodiment of the present invention, including:

[0063] S1. After determining that the user selects to perform dynamic processing of the language model on a first file in a preset modality, the first file is subjected to content segmentation processing to generate a plurality of content regions, and the first file is an image.

[0064] When the user selects to perform dynamic processing of the language model on a first file in a preset modality, the system will quickly start the content segmentation process. This operation is the basis for subsequent precise interaction and language result output. By dividing the image into multiple regions according to the content, it provides a more targeted information acquisition way for the user.

[0065] In some embodiments, the step of subjecting the first file to content segmentation processing to generate a plurality of content regions includes:

[0066] S11. Perform graphic comparison on the first file based on a preset element database to determine first image elements in the first file that pass the comparison.

[0067] The system will import the first file and perform in-depth graphic comparison with the preset element database. The preset element database covers a large number of common image element templates, from desks, chairs, and electrical appliances in daily life, to trees, flowers, and plants in natural scenes, to various geometric figures, etc. During the comparison process, the system uses advanced image recognition algorithms to match each part of the image with the templates in the database one by one. For example, for an image of a "park scene", the system successfully identifies first image elements such as "bench", "tree", and "flower" after comparison. This graphic comparison technology can quickly and accurately locate various elements existing in the image, providing a key basis for subsequent coordinate processing and content region determination.

[0068] S12. After coordinate processing of the first file, obtain the coordinate intervals of each first image element, and generate content coordinate intervals of the first image elements after processing the coordinate intervals.

[0069] After completing the recognition of image elements, the system performs coordinate processing on the first file. In a digital image, each pixel corresponds to a unique horizontal and vertical coordinate. By establishing a coordinate system for the entire image, the coordinate range of each first image element can be accurately obtained. Taking the image element "bench" as an example, assume that after coordinate processing, its horizontal coordinate range is [100, 200], and its vertical coordinate range is [150, 250].

[0070] Among them, the content coordinate range of the first image element generated after processing the coordinate range includes:

[0071] S121, obtain the extreme values of the horizontal coordinates and the extreme values of the vertical coordinates of the coordinate range.

[0072] To ensure that all information of the image element can be fully covered and avoid partial information loss caused by boundary truncation, the system will further process the obtained coordinate range to generate the content coordinate range.

[0073] S122, increase the preset value to the maximum value of the extreme values of the horizontal coordinates and the maximum value of the extreme values of the vertical coordinates respectively, and decrease the preset value to the minimum value of the extreme values of the vertical coordinates and the minimum value of the extreme values of the vertical coordinates respectively, to obtain the content coordinate range.

[0074] Exemplarily, the system will increase the preset value (assumed to be 10) to the maximum value of the extreme values of the horizontal coordinates (here it is 200) to get 210; decrease the preset value (10) to the minimum value of the horizontal coordinates (100) to get 90. Similarly, increase the preset value (10) to the maximum value of the extreme values of the vertical coordinates (250) to get 260; decrease the preset value (10) to the minimum value of the vertical coordinates (150) to get 140. Finally, the content coordinate range of the "bench" is the horizontal coordinate range of [90, 210] and the vertical coordinate range of [140, 260].

[0075] S13, regard all pixel points within the content coordinate range as the corresponding content area.

[0076] The content coordinate range processed through the above steps can more comprehensively contain the relevant pixel points of the image element, providing a more reliable range definition for determining the content area in the subsequent process.

[0077] S2, generate a path record layer corresponding to the first file.

[0078] In the language model dynamic processing method of the present invention, generating a path record layer corresponding to the first file is a crucial step, providing an indispensable basis for determining dynamic interaction content based on user interaction behaviors later. By recording the pixel point information triggered during the user's interaction with the image, the system can deeply analyze the user's operation trajectory, thereby generating accurate and personalized language results.

[0079] In some embodiments, generating the path record layer corresponding to the first file includes:

[0080] S21, obtaining the coordinates of each pixel point in the first file.

[0081] Within the scope of digital images, each image is composed of a large number of pixel points, which are the basic carriers of image information. The system will traverse each pixel point in the first file (i.e., the image), and using mature image parsing techniques, accurately obtain the horizontal and vertical coordinate values of each pixel point in the image coordinate system. Taking a common color image with a resolution of 1920×1080 as an example, the system will start from the first pixel point in the upper left corner of the image, and successively read its horizontal coordinate as 0 and vertical coordinate as 0; then read the horizontal coordinate of its adjacent pixel on the right as 1 and vertical coordinate as 0, and so on, until traversing the last pixel point in the lower right corner of the image, that is, the horizontal coordinate is 1919 and the vertical coordinate is 1079. Through this comprehensive and detailed traversal method, the system can completely master the position information of each pixel point in the image.

[0082] S22, establishing layer coordinates corresponding to the coordinates of each pixel point to generate a path record layer.

[0083] After successfully obtaining the coordinates of each pixel point in the first file, the system will establish a corresponding layer coordinate system for these coordinates. This layer coordinate system is interrelated and one-to-one corresponding to the coordinate system of the image itself. In essence, in an independent layer structure, a mapping record is created for the coordinates of each pixel point. This path record layer is like a transparent "overlay", closely adhering to the original image, always ready to record every interaction operation between the user and the image, providing key data support for subsequent in-depth analysis of user behaviors and determination of dynamic interaction content.

[0084] S3, after determining that the user triggers any pixel point in the path record layer, determining dynamic interaction content based on the spatio-temporal relationship of the triggered pixel point.

[0085] In the language model dynamic processing system of the present invention, step S3 is in the core interaction analysis link, and it undertakes the key task of converting the user's interaction operation with the path record layer into dynamic interaction content that the system can understand.

[0086] In some embodiments, after determining that any pixel point in the path recording layer is triggered by the user, determining dynamic interaction content based on the spatio-temporal relationship of the triggered pixel point includes:

[0087] S31, obtaining the first pixel point in the triggered path recording layer in real time, and extracting all the second pixel points triggered historically within a preset time period.

[0088] The system has the ability to monitor the path recording layer in real time. Once a user trigger operation is detected, it will quickly capture the first pixel point in the triggered path recording layer. For example, when the user clicks on a certain position on the image interface, the pixel point in the corresponding path recording layer at that position is the first pixel point. Suppose its coordinates are (150, 200). At the same time, the system will trace back all the second pixel points triggered historically within a preset time period. The preset time period can be flexibly set according to the actual application scenario and user interaction habits. For example, it is commonly set to the past 5 seconds or 10 seconds, etc. Within this time period, the system will retrieve and extract all the pixel point information triggered by the user before. For example, within the past 5 seconds, the user has also triggered pixel points with coordinates (120, 180), (130, 190), etc. These pixel points are the second pixel points. In this way, the system comprehensively collects the key information of the user's interaction with the image within a period of time.

[0089] S32, generating a pixel point sequence based on the time sequence of the first pixel point and the second pixel point, and sorting the first pixel point and the second pixel point in ascending order of time based on the pixel point sequence, so that the first pixel point and the second pixel point respectively have corresponding sorting numbers.

[0090] Based on the obtained first pixel point and second pixel point, the system will generate a pixel point sequence according to their trigger time sequence. The system arranges these pixel points in chronological order to form an ordered sequence. Then, the system will further process this pixel point sequence, sorting the first pixel point and the second pixel point in ascending order of time. In this way, each pixel point is assigned a corresponding sorting number, clearly showing the user's operation sequence. For example, the latest triggered first pixel point (150, 200) has a sorting number of 1, while the previously triggered (130, 190) has a sorting number of 2, (120, 180) has a sorting number of 3, etc. This sorting method provides a clear order basis for subsequent trace processing and instruction recognition, enabling the system to accurately restore the user's interaction process.

[0091] S33, performing trace processing on the first pixel point and the second pixel point in order based on the sorting number to determine the dynamic interaction content.

[0092] Among them, the sequential tracing process of the first pixel point and the second pixel point based on the sorting number to determine the dynamic interaction content includes:

[0093] S331. Connect the first pixel point and the second pixel point in sequence based on the sorting number to generate a tracing route.

[0094] The system connects the first pixel point and the second pixel point in sequence based on the sorting number to generate a tracing route. Taking the previous pixel points as an example, the system starts from the pixel point with the sorting number 1 (150, 200), and connects the pixel points with the sorting numbers 2 (130, 190) and 3 (120, 180) in sequence to form a continuous trajectory. This tracing route intuitively reflects the operation trajectory of the user on the image.

[0095] S332. Extract the graph formed by the tracing route and perform a graph comparison with a preset instruction database to determine the first instruction that passes the comparison. Each instruction has a preset response strategy.

[0096] After generating the tracing route, the system extracts the graph features formed by the tracing route and performs a graph comparison with the preset instruction database. The preset instruction database stores various common user operation instruction graph templates and their corresponding instruction meanings in advance. For example, a circle may correspond to the "content recognition" instruction, and a straight line may correspond to the "positioning" instruction, etc. The system uses an advanced graph matching algorithm to compare the tracing route graph with the templates in the database one by one to determine the first instruction that passes the comparison. In this way, the system can accurately identify the operation intention of the user and provide a direction for subsequent responses.

[0097] S333. Extract the dynamic interaction content corresponding to this instruction based on the first instruction.

[0098] This solution will analyze in combination with the first instruction to obtain the corresponding dynamic interaction content.

[0099] Among them, the extraction of the dynamic interaction content corresponding to this instruction based on the first instruction includes:

[0100] S3331. If the preset response strategy of the first instruction is a positioning strategy, extract the abscissa extreme values and ordinate extreme values of all the first pixel points and the second pixel points in the tracing route.

[0101] When the system determines that the preset response strategy for the first instruction is the positioning strategy, it will first conduct a detailed analysis of all the first pixel points and second pixel points involved in the tracing route. The system will traverse these pixel points and accurately extract the extreme values of their abscissas and ordinates. For example, assuming the tracing route is a "plus" sign, the system will find the minimum and maximum values of the abscissa and the minimum and maximum values of the ordinate. These extreme value coordinates will serve as the key data basis for determining the positioning midpoint subsequently.

[0102] S3332. Determine the abscissa midpoint and the ordinate midpoint based on the extreme values of the abscissa and the ordinate to obtain the positioning midpoint.

[0103] Based on the extreme values of the abscissa and the ordinate extracted in the previous step, the system will perform precise calculations. For the abscissa, add the maximum value and the minimum value and then divide the sum by 2 to obtain the abscissa midpoint. Similarly, for the ordinate, obtain the ordinate midpoint. The coordinate point formed by combining these two midpoints is the positioning midpoint. The positioning midpoint represents the central position of the user's operation trajectory and provides a core reference for determining the area where the user is located subsequently.

[0104] S3333. Determine the initial segmentation area of the path recording layer where the positioning midpoint is located, and output the corresponding output language result according to the initial segmentation area where the positioning midpoint is located.

[0105] Among them, the initial segmentation area is, for example, dividing the path recording layer into multiple areas, resulting in including multiple initial segmentation areas. This solution will output the corresponding output language result in combination with the initial segmentation area where the positioning midpoint is located.

[0106] It is worth mentioning that for blind users, obtaining image information has always been a difficult problem. With the dynamic processing method of the language model of the present invention, blind users can interact with images by touching the screen and other means. In the field of children's education, children have a strong demand for understanding and exploring images. Taking an image of a fairy tale scene as an example, children use a device equipped with this system to view the image. This solution is not limited to the above examples.

[0107] See Figure 2 , in some embodiments, after obtaining the tracing route, this solution can also perform the following processing steps, including:

[0108] Judge whether the triggering between the first pixel point and the second pixel point in the triggering path recording layer is continuous triggering or segmented triggering. Judge whether the triggering between the first pixel point and the second pixel point in the triggering path recording layer is continuous triggering or segmented triggering. The system makes a judgment by analyzing the spatial position relationship of the pixel points in the path recording layer. If the spatial positions of adjacent pixel points are adjacent (for example, within a certain pixel distance threshold), it is determined as continuous triggering; otherwise, it is determined as segmented triggering.

[0109] If it is a continuous trigger, use the tracing route as the finally determined tracing route. If it is a continuous trigger, use the tracing route as the finally determined tracing route. This means that the user's operation is coherent and no further processing of the path is required.

[0110] If it is a segmented trigger, determine each segmented trigger path segment, sort them according to the generation order of the segmented trigger path segments to obtain a path segment sequence. If it is a segmented trigger, determine each segmented trigger path segment, sort them according to the generation order of the segmented trigger path segments to obtain a path segment sequence. For example, if the user first draws a straight line and then pauses for a while and draws another straight line, the system will identify these two segmented trigger path segments and sort them according to the time of generation.

[0111] Group adjacent segmented trigger path segments in the path segment sequence in pairs, determine the last pixel point of the previous segmented trigger path segment in each group as the detachment pixel point, and determine the first pixel point of the subsequent segmented trigger path segment in each group as the starting pixel point. Group adjacent segmented trigger path segments in the path segment sequence in pairs, determine the last pixel point of the previous segmented trigger path segment in each group as the detachment pixel point, and determine the first pixel point of the subsequent segmented trigger path segment in each group as the starting pixel point.

[0112] Update the route between the detachment pixel point and the starting pixel point in the corresponding group in the tracing route to a dotted line or delete it to obtain the finally determined tracing route. Update the route between the detachment pixel point and the starting pixel point in the corresponding group in the tracing route to a dotted line or delete it to obtain the finally determined tracing route. Through this processing method, the user's effective operation path can be presented more clearly, avoiding path confusion caused by pauses or incorrect operations during the user's operation process, which may affect subsequent instruction recognition and language result output.

[0113] In the above embodiment, the initial segmentation area of the path record layer is obtained through the following steps, including:

[0114] Obtain the number of horizontal pixel points and the number of vertical pixel points in the path record layer. To determine the initial segmentation area of the path record layer where the positioning midpoint is located, the system first obtains the number of horizontal pixel points and the number of vertical pixel points in the path record layer. Assume that the number of horizontal pixel points in the path record layer is 800 and the number of vertical pixel points is 600.

[0115] The number of horizontal pixel points is equally divided according to a preset horizontal equal division value to obtain a plurality of horizontal equal division points, and the number of vertical pixel points is equally divided according to a preset vertical equal division value to obtain a plurality of vertical equal division points. Exemplarily, the preset horizontal equal division value is set to 8, and the preset vertical equal division value is set to 6. Then, the number of horizontal pixel points is equally divided, 800÷8 = 100, to obtain 8 horizontal equal division points, with a spacing of 100 pixels between each equal division point; the number of vertical pixel points is equally divided, 600÷6 = 100, to obtain 6 vertical equal division points, with a spacing of 100 pixels between each equal division point.

[0116] Based on the horizontal equal division points and vertical equal division points, generate equal division lines to equally divide the path recording layer, obtain the initial segmentation regions of the path recording layer, and add corresponding language orientation information to each initial segmentation region. The output language result includes the language orientation information.

[0117] Based on these horizontal equal division points and vertical equal division points, the system generates equal division lines and divides the path recording layer into 48 initial segmentation regions. At the same time, the system will add corresponding language orientation information to each initial segmentation region. For example, the upper left corner region can be marked as "upper left corner region", the upper right corner region can be marked as "upper right corner region", etc. By judging the coordinates of the positioning midpoint (115, 165), it can be known that it is located in the 2nd horizontal segmentation region from left to right and the 2nd vertical segmentation region from top to bottom, and the corresponding language orientation information is "upper left middle region". The system outputs the corresponding output language result according to this information, such as "The position you operate is in the upper left middle region of the image".

[0118] In some other embodiments, the extracting of the dynamic interaction content corresponding to the current instruction based on the first instruction includes:

[0119] If the preset response strategy of the first instruction is the content recognition strategy, then extract the abscissa extreme values and ordinate extreme values of all the first pixel points and second pixel points in the tracing route.

[0120] If the preset response strategy of the first instruction is the content recognition strategy, the system will also extract the abscissa extreme values and ordinate extreme values of all the first pixel points and second pixel points in the tracing route.

[0121] Generate a selected region based on the abscissa extreme values and ordinate extreme values.

[0122] For example, assume that the abscissa extreme values determined from the pixel point coordinates on the tracing route are [90, 140], and the ordinate extreme values are [130, 170]. Based on these extreme values, the system generates a selected region, and the range of this selected region on the path recording layer is a rectangular region with an abscissa from 90 to 140 and an ordinate from 130 to 170.

[0123] Map the pixel points within the selected area to the first file to determine the dynamic interaction content.

[0124] The system maps the pixel points within the generated selected area to the first file (i.e., the original image). Since there is a corresponding relationship between the coordinate of the path recording layer and the original image, through this mapping, the system can determine the part of the original image corresponding to the selected area. For example, in the "park scene" image, the selected area may cover the bench and the surrounding grass. The system takes this part of the image content as the dynamic interaction content, providing a basis for determining the output language result subsequently, such as analyzing the image elements within this area, determining the intersection with the "bench" image element, and further determining the final output language result according to the intersection situation.

[0125] S4. Determine at least one content or at least one content area output language result based on the dynamic interaction content.

[0126] In the dynamic processing flow of the language model of the present invention, step S4 is a key link to convert the previously determined dynamic interaction content into a language result understandable by the end user. Through the detailed analysis of the selected area, image elements, and content areas, the system can accurately output a language description that conforms to the user's interaction intention, realizing the key conversion from image interaction operations to natural language feedback.

[0127] In some embodiments, the determining at least one content or at least one content area output language result based on the dynamic interaction content includes:

[0128] S41. Obtain the first image element whose pixel point coordinate position intersects with the selected area, determine the number of pixel points in the intersection to obtain the first intersection number, and count the first total number of pixel points within the first image element.

[0129] When the system determines that a language result needs to be output based on dynamic interactive content, it first focuses on the selected area. In the scenario based on the content recognition strategy, the system will obtain the first image element whose pixel point coordinate position intersects with the selected area. For example, in the "park scene" image, if the coordinate range of the selected area is [110, 160] for the abscissa and [170, 210] for the ordinate, the system will use the coordinate matching algorithm to retrieve the first image element that has been recognized in the preset element database and find that the "bench" image element intersects with the selected area. Then, the system will accurately count the number of pixel points in the intersection to obtain the first intersection quantity. Suppose it is determined through traversal calculation that the number of pixel points in this intersection is 600. At the same time, the system will count the first total quantity of the pixel points within the first image element of "bench", which is supposed to be 1000. These quantity statistics provide a quantitative basis for subsequent ratio calculation and language result judgment.

[0130] S42. Calculate the ratio of the first intersection quantity to the first total quantity to obtain the first ratio. If the first ratio is greater than the preset value, then use the content of the corresponding first image element as the output language result.

[0131] The system will calculate the ratio of the first intersection quantity to the first total quantity, that is, 600÷1000 = 0.6, to obtain the first ratio. The preset value can be flexibly set according to the actual application scenario and requirements. For example, it is set to 0.5. When the first ratio is greater than the preset value, it means that the degree of coincidence between the selected area and this first image element is relatively high, and it can more accurately reflect the user's attention to this image element. At this time, the system will use the content of the corresponding first image element as the output language result. Taking the "bench" image element as an example, the system may output "This is a brown wooden bench that can accommodate two people to rest", such a language description accurately reflects the information of the image element that the user is concerned about.

[0132] S43. If it is determined that all first ratios are less than or equal to the preset value, then determine a content area to output the language result.

[0133] If it is determined that all first ratios are less than or equal to the preset value, it means that the degree of coincidence between the selected area and a single image element is not high enough to clearly point to a specific image element.

[0134] Among them, the determination of a content area to output the language result includes:

[0135] S431. Obtain the content area whose pixel point coordinate position intersects with the selected area, and determine the number of pixel points in the intersection to obtain the second intersection quantity.

[0136] The system will change its strategy to obtain the content areas that intersect with the pixel coordinates of the selected area. For example, in the "park scene" image, in addition to the "bench" image element, there may also be a content area containing "flower beds" that intersects with the selected area. The system will determine the number of pixels in these intersections to obtain the second intersection quantity.

[0137] S432. If it is determined that there is only one content area with an intersection with the selected area, the language result of the corresponding content area will be output.

[0138] If it is determined that there is only one content area with an intersection with the selected area, at this time, the system will output the language result of the corresponding content area. For example, if the content area contains a "bench", the system may output "This area contains a bench".

[0139] S433. If it is determined that there are multiple content areas with intersections with the selected area, the content area with the larger second intersection quantity will be selected to output the language result, and the output language result includes all the content within the content area.

[0140] If it is determined that there are multiple content areas with intersections with the selected area, the system will further compare the intersection degree between these content areas and the selected area.

[0141] By comparing the second intersection quantities, the content area with the larger second intersection quantity is selected to output the language result. Suppose there is a content area containing a "bench" with a second intersection quantity of 400, and another content area containing "flowers" with a second intersection quantity of 300. The system will select the content area containing the "bench" with the larger second intersection quantity to output the language result, and the output language result will include all the content within the content area, such as "This area contains a bench", so as to provide language feedback to the user.

[0142] See Figure 3 , which is a schematic structural diagram of a language model dynamic processing system provided by an embodiment of the present invention. The system includes:

[0143] A judgment module, configured to, after judging that the user selects to perform language model dynamic processing on a first file of a preset modality, perform content segmentation processing on the first file to generate multiple content areas, where the first file is an image;

[0144] A generation module, configured to generate a path record layer corresponding to the first file;

[0145] A determination module, configured to, after judging that any pixel point in the path record layer is triggered by the user, determine dynamic interaction content based on the spatio-temporal relationship of the triggered pixel point;

[0146] An output module, configured to determine at least one content or at least one content area output language result based on the dynamic interaction content.

[0147] The present invention also provides a storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the methods provided by the above various embodiments.

[0148] Among them, the storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). In addition, the ASIC can be located in the user equipment. Of course, the processor and the storage medium can also exist as discrete components in the communication device. The storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0149] The present invention also provides a program product, which includes execution instructions stored in a storage medium. At least one processor of the device can read the execution instructions from the storage medium, and at least one processor executes the execution instructions to enable the device to implement the methods provided by the above various embodiments.

[0150] In the above embodiments of the terminal or the server, it should be understood that the processor can be a central processing unit (CPU for short), and can also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the present invention can be directly embodied as being completed by the execution of the hardware processor, or by the combination of the hardware and software modules in the processor.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A language model dynamic processing method, characterized in that: include: After determining that the user chooses to perform language model dynamic processing on a first file of a preset modality, the first file is segmented according to content to generate a plurality of content areas, wherein the first file is an image; Generate a path record layer corresponding to the first file; After determining that the user triggers any pixel point in the path record layer, the dynamic interaction content is determined based on the spatiotemporal relationship of the triggered pixel point, including: Get the pixel points in the trigger path record layer and generate the tracking route; Extracting the graph formed by the tracing route and comparing it with the preset instruction database, determining the first instruction that passes the comparison, each instruction has a preset response strategy, and the preset response strategy includes a positioning strategy and a content identification strategy; Extracting corresponding dynamic interactive content based on the first instruction; At least one content or at least one content area outputs a language result based on the dynamic interactive content.

2. The method according to claim 1, characterized in that The step of segmenting the first file according to the content to generate a plurality of content areas includes: Performing graphic comparison on the first file based on a preset element database to determine a first image element in the first file that passes the comparison; After coordinate processing of the first file, a coordinate interval of each first image element is obtained, and after processing the coordinate interval, a content coordinate interval of the first image element is generated; All pixels within the content coordinate interval are regarded as the corresponding content area.

3. The method according to claim 2, characterized in that The processing of the coordinate interval to generate a content coordinate interval of the first image element includes: Get the horizontal and vertical extreme values ​​of the coordinate interval; The maximum value among the extreme values ​​of the horizontal coordinate and the maximum value among the extreme values ​​of the vertical coordinate are increased by a preset value, and the minimum value among the extreme values ​​of the vertical coordinate and the minimum value among the extreme values ​​of the vertical coordinate are decreased by a preset value to obtain a content coordinate interval.

4. The method according to claim 2, characterized in that: The generating of the path record layer corresponding to the first file includes: Get the coordinates of each pixel in the first file; Establish the layer coordinates corresponding to the coordinates of each pixel point and generate a path record layer.

5. The method according to claim 2, characterized in that: After determining that a user triggers any pixel point in the path record layer, determining the dynamic interaction content based on the spatiotemporal relationship of the triggered pixel point includes: Acquire the first pixel point in the trigger path record layer in real time, and extract the second pixel points of all historical triggers within a preset time period; Generate a pixel point sequence based on the time sequence of the first pixel point and the second pixel point, and sort the first pixel point and the second pixel point from near to far in time based on the pixel point sequence, so that the first pixel point and the second pixel point have corresponding sorting numbers respectively; The first pixel point and the second pixel point are traced in sequence based on the sequence number to determine the dynamic interactive content.

6. The method according to claim 5, characterized in that The step of tracking the first pixel point and the second pixel point in sequence based on the sequence number to determine the dynamic interactive content includes: The first pixel point and the second pixel point are connected in sequence based on the sequence number to generate a tracing route.

7. The method according to claim 6, characterized in that The extracting the dynamic interactive content corresponding to the current instruction based on the first instruction includes: If the preset response strategy of the first instruction is a positioning strategy, extracting the horizontal coordinate extreme values ​​and the vertical coordinate extreme values ​​of all the first pixel points and the second pixel points in the tracing route; Determine the middle value of the abscissa and the middle value of the ordinate based on the extreme value of the abscissa and the extreme value of the ordinate to obtain a positioning middle point; Determine the initial segmentation area of ​​the path record layer where the positioning intermediate point is located, and output the corresponding output language result according to the initial segmentation area where the positioning intermediate point is located.

8. The method according to claim 7, characterized in that The initial segmentation area of ​​the path record layer is obtained by the following steps, including: Get the number of horizontal and vertical pixels in the path record layer; The number of horizontal pixel points is equally divided according to a preset horizontal equal division value to obtain a plurality of horizontal equally divided points, and the number of vertical pixel points is equally divided according to a preset vertical equal division value to obtain a plurality of vertical equally divided points; Based on the horizontal and vertical dividing points, the dividing lines are generated to divide the path record layer into equal parts to obtain the initial segmentation area of ​​the path record layer, and corresponding language orientation information is added to each initial segmentation area. The output language result includes the language orientation information.

9. The method according to claim 6, characterized in that The extracting the dynamic interactive content corresponding to the current instruction based on the first instruction includes: If the preset response strategy of the first instruction is a content recognition strategy, extracting the horizontal coordinate extreme values ​​and the vertical coordinate extreme values ​​of all the first pixel points and the second pixel points in the tracing route; Generate a selected area based on the abscissa extreme value and the ordinate extreme value; The pixels in the selected area are mapped to the first file to determine the dynamic interactive content.

10. The method according to claim 9, characterized in that The step of determining at least one content or at least one content area output language result based on the dynamic interactive content comprises: Acquire a first image element that intersects with the coordinate position of the pixel point in the selected area, determine the number of the pixel points of the intersection to obtain a first intersection number, and count a first total number of the pixel points in the first image element; Calculating a ratio of the first intersection quantity to the first total quantity to obtain a first proportion, and if the first proportion is greater than a preset value, using the content of the corresponding first image element as the output language result; If it is determined that all first proportions are less than or equal to a preset value, a content area is determined to output a language result.

11. The method according to claim 10, characterized in that The step of determining a content area output language result includes: Obtaining a content area that intersects with the pixel coordinate positions of the selected area, and determining the number of pixels of the intersection to obtain a second intersection number; If it is determined that there is only one content area that has an intersection with the selected area, the corresponding content area will be determined to output the language result; If it is determined that there are multiple content areas that have intersections with the selected area, then the content area with the second largest number of intersections is selected to output the language result, and the output language result includes all the content in the content area.

12. A language model dynamic processing system according to any one of claims 1 to 11, characterized in that: a judgment module, configured to generate a plurality of content regions by segmenting the first file according to the content after judging that the user chooses to perform language model dynamic processing on the first file of the preset modality, wherein the first file is an image; A generation module, used for generating a path record layer corresponding to the first file; A determination module, for determining the dynamic interaction content based on the spatiotemporal relationship of the triggered pixel points after determining that the user has triggered any pixel point in the path record layer; The output module is used to determine at least one content or at least one content area based on the dynamic interactive content and output a language result.

13. A storage medium, characterized in that The storage medium stores a computer program, which is used to implement the method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Intelligent touch reading method and touch reading device

    CN109255989A

  • Image recognition method and device, equipment and storage medium

    CN112200167A