6  sillysplines: Data generation using linear splines

After addressing the XAI explanation layer and the visualization of the dataset layer, the next step we addressed is the synthetic data generation for model testing.

In any two-dimensional area, there are numerous ways to define a boundary for a classification task. While complex, disjointed boundaries can exist, a single, continuous line is often preferred when teaching machine learning concepts. As a pedagogical tool, it’s far more intuitive for students and end users to visually separate two regions with a continuous line, making it easier to conceptualize how a model is attempting to learn and generalize.

The initial design of the tool

However, to ensure that the generated boundary provides a suitable challenge for predictive models, the default design is intentionally complex. Instead of a simple line parallel to an axis, the design consists of a mix of oblique and axis-parallel segments. This approach ensures that the downstream classification task is not trivial and compels the model to learn non-linear relationships. Furthermore, this mixed design helps the model developer discern which of the two features is more influential in the model’s final decision-making process.

Key points of consideration

When developing this tool several technical considerations had to be taken to translate the user input into a usable decision boundary.

A dual coordinate system was implemented to translate between the screen coordinates and the data coordinates. The display coordinates range from 0 to 640 pixels, while the data coordinates are normalized to a -10 to 10 range. This separation is crucial because it allows the visualization to be resolution-independent and makes the data more meaningful for mathematical operations or external processing. The transformation also includes a y-axis flip (subtracting from the maximum value) because SVG coordinates have their origin at the top-left, while most mathematical coordinate systems place the origin at the bottom-left. When implementing this pattern, care was taken to be mindful of the order of operations during the coordinate conversion, especially when dealing with the flipped y-axis.

When dragging begins on an existing point, it selects that point for manipulation. However, when dragging starts on empty space, the system dynamically creates a new point at that location and immediately begins dragging it. This creates an intuitive user experience where clicking anywhere adds a point.

The visualization creates two complementary polygons - one above the draggable line and one below it. The upper polygon is constructed by connecting the top corners of the canvas with all the user-defined points, while the lower polygon connects the bottom corners with the same points. Both polygons require coordinate sorting to ensure proper edge connections, but this is done at render time (ie. when drawing the polygons on the interface) rather than modifying the original data structure. A key assumption is that the polygon rendering assumes points are meant to be connected in x-coordinate order, which works well for function-like curves.