LUMIERA.clone

Author	SHA1	Message	Date
Ichthyostega	54f238c510	Scheduler-test: further configuration options for flexible testing The next goal is to determine basic performance characteristics of the Scheduler implementation written thus far; to help with these investigations some added flexibility seems expedient - the ability to define a per-job base expense - added flexibility regarding the scheduling of dependencies This changeset introduces configuration options	2024-01-06 01:32:14 +01:00
Ichthyostega	032e4f6db5	Scheduler-test: extract search algo into lib	2024-01-06 01:23:00 +01:00
Ichthyostega	e52aed0b3c	Scheduler-test: simplify binary search implementation While the idea with capturing observation values is nice, it definitively does not belong into a library impl of the search algorithm, because this is usage specific and grossly complicates the invocation. Rather, observation data can be captured by side-effect from the probe-λ holding the actual measurement run.	2024-01-04 02:37:05 +01:00
Ichthyostega	bdc1b089d7	Scheduler-test: binary search working - Result found in typically 6-7 steps; - running 20 instead of 30 samples seems sufficient Breaking point in this example at stress-Factor 0.47 with run-time 39ms	2024-01-04 01:37:43 +01:00
Ichthyostega	bf1eac02dd	Scheduler-test: binary search over continuous domain - textbook implementation - capture results from visited points - average results form the last three points to damp statistic fluctuations	2024-01-04 00:04:22 +01:00
Ichthyostega	54e489b9b6	Scheduler-test: define criterion for breaking point This statistical criterion defines when to count observed Scheduler performance as loosing control. The test is comprised of three observations, which all must be confirmed: - an individual run counts as accidentally failed when the execution slips away by more than 2ms with respect to the defined overall schedule. When more than 55% of all observed runs are considered as failed, the first condition is met - moreover, the observed standard derivation must also surpass the same limit of > 2ms, which indicates that the Scheduling mechanism is under substantial strain (on average, the slip is ~ 200µs) - the third condition is that the ''averaged delta'' has surpassed 4ms, which is 2 times the basic failure indicator. These conditions are based on watching the Scheduler in operation; typically all three conditions slip away by large margin after a very narrow yet critical increase in the stress level. Using three conditions together should improve robustness; often the problems to keep up with the schedule build up over some parameter range, yet the actual decision should be based on complete loss of control.	2024-01-03 21:11:20 +01:00
Ichthyostega	fda34a42ca	Scheduler-test: incorporate statistics computation adapt the code written yesterday explicitly for the test case into the new framework for performing a stress-test run. Notable difference: times converted to millisecond immediately	2024-01-03 16:27:07 +01:00
Ichthyostega	e704f4aae0	Scheduler-test: build configurable measurement setup Elaborate the draft to include all the elements used directly in the test case thus far; the goal is to introduce some structuring and leave room for flexible confguration, while implementing the actual binary search as library function over Lambdas. My expectation is to write a series of individual test instances with varying parameters; while it seems possible to add further performance test variations into that scheme later on.	2024-01-03 02:18:15 +01:00
Ichthyostega	6f4bd150fd	Scheduler-test: draft a structure to formalise these investigations - the goal is to run a binary search - the search condition should be factored out - thus some kind of framework or DSL is required, to separate the technicalities of the measurement from the specifics of the actual test case.	2024-01-03 00:45:17 +01:00
Ichthyostega	29699991a0	Scheduler-test: watch statistics with increasing stress - repeated invocations of the same test setup for statistics - the usual nasty 64-node graph with massive fork out - limit concurrency to 4 cores - tabulate data to look for clues regarding a trigger criteria Hypothesis: The Scheduler slips off schedule when all of the following three criteria are met: - more than 55% glitches with Δ > 2ms - σ > 2ms - ∅Δ > 4ms	2024-01-02 18:44:20 +01:00
Ichthyostega	a56afbaf62	Scheduler-test: bugfix in computation-load ...this one was quite silly: obviously we need a separate instance of the memory block ''per invocation'', otherwise concurrent invocations would corrupt each other's allocation. The whole point of this variant of the computation-load is to access a ''private'' memory block...	2024-01-02 16:24:24 +01:00
Ichthyostega	0c9485294e	Scheduler-test: experiment with varying stress levels	2024-01-02 15:45:40 +01:00
Ichthyostega	bb2bbc0e02	Scheduler-test: verify adapted schedule with stress-factor - schedule can now be adapted to concurrency and expected distribution of runtimes - additional stress factor to press the schedule (1.0 is nominal speed) - observed run-time now without Scheduler start-up and pre-roll - document and verify computed numbers	2024-01-02 00:24:28 +01:00
Ichthyostega	813f8721f7	Scheduler-test: build adapted schedule ...based on the adapted time-factor sequence implemented yesterday in TestChainLoad itself - in this case, the TimeBase from the computation load is used as level speed - this »base beat« is then modulated by the timing factor sequence - working in an additional stress factor to press the schedule uniformly - actual start time will be added as offset once the actual test commences	2024-01-01 22:48:27 +01:00
Ichthyostega	f4dd309476	Scheduler-test: change test setup to use a schedule table ...up to now, we've relied on a regular schedule governed solely by the progression of node levels, with a fixed level speed defaulting to 1ms per level. But in preparation of stress tesging, we need a schedule adapted to the expected distribution of computation times, otherwise we'll not be able to factor out the actual computation graph connectivity. The goal is to establish a distinctive breaking point when the scheduler is unable to cope with the provided schedule.	2024-01-01 21:07:16 +01:00
Ichthyostega	55cb028abf	Scheduler-test: document and verify weight adapted timing The helper developed thus far produces a sequence of weight factors per level, which could then be multiplied with an actual delay base time to produce a concrete schedule. These calculations, while simple, are difficult to understand; recommended to use the values tabulated in this test together with a `graphviz` rendering of the node graph (🠲 `printTopologyDOT()`)	2023-12-31 21:59:41 +01:00
Ichthyostega	e9e7d954b1	Scheduler-test: formula to guess expense The intention is to establish a theoretical limit for the expense, given some degree of concurrency. In reality, the expense should always be greater, since the time is not just split by the number of cores; rather we need to chain up existing jobs of various weight on the available cores (which is a special case of the box packing problem). With this formula, an ideal weight factor can be determined for each level, and then summing up the sequence of levels gives us a guess for a sensible timing for the overall scheduler	2023-12-31 03:14:59 +01:00
Ichthyostega	409a60238a	Scheduler-test: extract a generic grouping iterator ...so IterExplorer got yet another processing layer, which uses the grouping mechanics developed yesterday, but is freely configurable through λ-Functions. At actual usage sit in TestChainLoad, now only the actual aggregation computation must be supplied, and follow-up computations can now be chained up easily as further transformation layers.	2023-12-31 00:41:01 +01:00
Ichthyostega	fec117039e	Scheduler-test: need this group aggregation as pipeline rather Yesterday I've written a simple loop-based implementation of a grouping aggregation to count the node weights per level. Unfortunately it turns out we'll use several flavours of this and we'd have to chain up postprocessing -- thus from a usage perspective it would be better to have the same functionality packaged as interator pipeline. This turns out to be surprisingly tricky and there is no suitable library function available, which means I'll have to write one myself. This changeset is the first step into this direction: reformulate the simple for-loop into a demand-driven grouping iterator	2023-12-30 02:15:38 +01:00
Ichthyostega	f04035a030	Scheduler-test: draft calculation of level-weight based schedule ...the idea is to use the sum of node weights per level to create a schedule, which more closely reflects the distribution of actual computation time. Hopefully such a schedule can then be squeezed or stretched by a time factor to find out a ''breaking point'', at which the Scheduler is no longer able to keep up.	2023-12-29 01:07:26 +01:00
Ichthyostega	47ae4f237c	Scheduler-test: investigate and fix further memory manager problem In-depth investigation and reasoning highlighted another problem, which could lead to memory corruption in rare cases; in the end I found a solution by caching the ''address'' of the current Epoch and re-validating this address on each Epoch-overflow. After some difficulties getting any reliable measurement for a Release-build, it turned out that this solution even ''improves performance by 22%'' Remark-1: the static blockFlow::Config prevents simple measurements by just recompiling one translation unit; it is necessary to build the relevant parts of Vault-layer with optimisation to get reliable numbers Remark-2: performing a full non-DEBUG build highlighted two missing header-inclusions to allow for the necessary template specialisations.	2023-12-28 02:13:24 +01:00
Ichthyostega	3716a5b3d4	Scheduler-test: address defects in memory manager ...discovered by during investigation of latest Scheduler failures. The root of the problems is that block overflow can potentially trigger expansion of the allocation pool. Under some circumstances, this on-the fly allocation requires a rotation of index slots, thereby invalidating existing iterators. While such behaviour is not uncommon with storage data structures (see std::vector), in this case it turns out problematic because due to performance considerations, a usage pattern emerged which exploits re-using existing storage »Slots« with known deadline. This optimisation seems to have significant leverage on the planning jobs, which happen to allocated and arrange a whole strike of Activities with similar deadlines. One of these problem situations can easily be fixed, since it is triggered through the iterator itself, using a delegate function to request a storage expansion, at which point the iterator is able to re-link and fix its internal index. This solution also has no tangible performance implications in optimised code. Unfortunately there remains one obscure corner case where such an pool expansion could also have invalidated other iterators, which are then used later to attach dependency relations; even a partial fix for that problem seems to cause considerable performance cost of about -14% in optimised code.	2023-12-27 00:16:03 +01:00
Ichthyostega	af680cdfd9	Scheduler-test: adapt tests to changed logic at entrance - now there can not be any direct dispatch anymore when entering events - thus there is no decision logic at entrance anymore - rather the work-function implementation moved down into Layer-2 - so add a unit-test like coverage there (integration in SchedulerService_test)	2023-12-27 00:16:03 +01:00
Ichthyostega	09f0e92ea3	Scheduler-test: reorganise planning-job entrance and coordination This amounts to a rather massive refactoring, prompted by the enduring problems observed when pressing the scheduler. All the various glitches and (fixed) crashes are related to the way how planning-jobs enter the schedule items, which is also closely tied to the difficulties getting the locking for planning-jobs correct. The solution pursued hereby is to reorder the main avenues into the scheduler implementation. There is now a streamlined main entrance, which always enqueues only, allowing to omit most checks and coordination. On the other hand, the complete coordination and dispatch of the work capacity is now shifted down into the SchedulerCommutator, thereby linking all coordination and access control close together into a single implementation facility. If this works out as intended - several repeated checks on the Grooming-Token could be omitted (performance) - the planning-job would no longer be able to loose / drop the Token, thereby running enforcedly single-threaded (as was the original intention) - since all planning effectively originates from planning-jobs, this would allow to omit many safety barriers and complexities at the scheduler entrance avenue, since now all entries just go into the queue. WIP: tests pass compiler, but must be adapted / reworked	2023-12-26 03:06:30 +01:00
Ichthyostega	dedfbf4984	Scheduler-test: investigate planning failure - fix mistake in schdule time for planning chunks (must use start, not end of chunk) - allow to configure the heuristics for pre-roll (time reserved for planning a node)	2023-12-23 21:38:51 +01:00
Ichthyostega	90ab20be61	Scheduler-test: press harder with long and massive graph ...observing multiple failures, which seem to be interconnected - the test-setup with the planning chunk pre-roll is insufficient - basically it is not possible to perform further concurrent planning, without getting behind the actual schedule; at least in the setup with DUMP print statements (which slowdown everything) - muliple chained re-entrant calls into the planning function can result - the ASSERTION in the Allocator was triggered again - the log+stacktrace indicate that there is still a Gap in the logic to protect the allocations via Grooming-Token	2023-12-22 00:33:51 +01:00
Ichthyostega	2cd51fa714	Scheduler-test: fix out-of-bound access ...causing the system to freeze due to excess memory allocation. Fortunately it turned out this was not an error in the Scheduler core or memory manager, but rather a sloppiness in the test scaffolding. However, this incident highlights that the memory manager lacks some sanity checks to prevent outright nonsensical allocation requests. Moreover it became clear again that the allocation happens ''already before'' entering the Scheduler — and thus the existing sanity check comes too late. Now I've used the same reasoning also for additional checks in the allocator, limiting the Epoch increment to 3000 and the total memory allocation to 8GiB Talking of Gibitbytes... indeed we could use a shorthand notation for that purpose...	2023-12-21 20:25:43 +01:00
Ichthyostega	707fbc2933	Scheduler-test: implement contention mitigation scheme while my basic assessment is still that contention will not play a significant role given the expected real world usage scenario — when testing with tighter schedule and rather short jobs (500µs), some phases of massive contention can be observed, leading to significant slow-down of the test. The major problem seems to be that extended phases of contention will effectively cause several workers to remain in an active spinning-loop for multiple microseconds, while also permanently reading the atomic lock. Thus an adaptive scheme is introduced: after some repeated contention events, workers now throttle down by themselves, with polling delays increased with exponential stepping up to 2ms. This turns out to be surprisingly effective and completely removes any observed delays in the test setup.	2023-12-20 20:25:17 +01:00
Ichthyostega	b497980522	Scheduler-test: guard memory allocations by grooming-token Turns out that we need to implemented fine grained and explicit handling logic to ensure that Activity planning only ever happens protected by the Grooming-Token. This is in accordance to the original design, which dictates that all management tasks must be done in »management mode«, which can only be entered by a single thread at a time. The underlying assumption is that the effort for management work is dwarfed in comparison to any media calculation work. However, in `5c6354882d` ...I discovered an insidious border condition, an in an attempt to fix it, I broke that fundamental assumpton. The problem arises from the fact that we do want to expose a public API of the Scheduler. Even while this is only used to ''seed'' a calculation stream, because any further planning- and management work will be performed by the workers themselves (this is a design decision, we do not employ a "scheduler thread") Anyway, since the Scheduler API ''is'' public, ''someone from the outside'' could invoke those functions, and — unaware of any Scheduler internals — will automatically acquire the Grooming-Token, yet never release it, leading to deadlock. So we need a dedicated solution, which is hereby implemented as a scoped guard: in the standard case, the caller is a management-job and thus already holds the token (and nothing must be done). But in the rare case of an »outsider«, this guard now ''transparently'' acquires the token (possibly with a blocking wait) and ''drops it when leaving scope''	2023-12-19 23:38:57 +01:00
Ichthyostega	f526360319	Scheduler-test: retract support for ''self-inhibition'' ...this feature seems to be no longer necessary now; leaving the actual implementation in-code for the time being, but removed it from the public access API.	2023-12-19 21:07:33 +01:00
Ichthyostega	c4807abf8a	Scheduler-test: planning for stress-tests	2023-12-19 21:06:23 +01:00
Ichthyostega	67036f45b0	Scheduler-test: Integration-test now running smoothly The last round of refactorings yielded significant improvements - parallelisation now works as expected - processing progresses closer to the schedule - run time was reduced The processing load for this test is tuned in a way to overload the scheduler massively at the end -- the result must be correct non the less. There was one notable glitch with an assertion failure from the memory manager. Hopefully I can reproduce this by pressing and overloading the Scheduler more...	2023-12-18 23:34:10 +01:00
Ichthyostega	ba82a446fd	Scheduler-test: address follow-up problem with depth-first The rework from yesterday turned out to be effective ... unfortunately a bit to much: since now late follow-up notifications take precedence, a single worker tends to process the complete chain depth-first, because the first chain will be followed and processed, even before the worker was able to post the tasks for the other branches. Thus this single worker is the only one to get a chance to proceed. After some consideration, I am now leaning towards a fundamental change, instead of just fixing some unfavourable behaviour pattern: while the language semantics remains the same, the scheduler should no longer directly dispatch into the next chain from λ-post. That is, whenever a POST / NOTIFY is issued from the Activity-chain, the scheduler goes through prioritisation. This has further ramifications: we do not need a self-inhibition mechanism any more (since now NOTIFY picks up the schedule time of the target). With these changes, processing seems to proceed more smoothly, albeit still with lots of contention on the Grooming token, at least in the example structure tested here.	2023-12-17 23:46:44 +01:00
Ichthyostega	1cbb6b7371	Scheduler-test: rework handling of notifications in the Activity-Language While the recent refactoring... `206c67cc` ...was a step into the right direction, it pushed too hard, overlooking the requirement to protect the scheduler contents and thus all of the Activity-chains against concurrent modification. Moreover, the recent solution still seems not quite orthogonal. Thus the handling of notifications was thoroughly reworked: - the explicit "double-dispatch" was removed, since actual usage of the language indicates that we only need notifications to Gate (and Hook), but not to any other conceivable Activity. - thus it seems unnecessary to turn "notification" into some kind of secondary work mode. Rather, it is folded as special case into the regular dispatch. This leads to new processing rules: - a POST goes into λ-post (obviously... that's its meaning) - a NOTIFY now passes its target into λ-post - λ-post invokes ''dispatch'' - and dispatching a Gate now implies to notify the Gate This greatly simplifies the »state machine« in the Activity-Language, but also incurs some limitations (which seems adequate, since it is now clear that we do not ''schedule'' or ''dispatch'' arbitrary Activities — rather we'll do this only with POST and NOTIFY, and all further processing happens by passing activation along the chain, without involving the Scheduler)	2023-12-16 23:47:50 +01:00
Ichthyostega	75b5eea2d3	Scheduler-test: option to require activation by scheduler use a feature of the Activity-Language prepared for this purpose: self-Inhibition of the Chain. This prevents a prerequisite-NOTIFY to trigger a complete chain of available tasks, before these tasks have actually reached their nominal scheduling time. This has the effect to align the computations much more strictly with the defined schedule	2023-12-14 01:49:46 +01:00
Ichthyostega	3e84224f74	Scheduler-test: force dependency-wait to wake-up job The main (test) thread is kept in a blocking wait until the planned schedule is completed. If however the schedule overruns, the wake-up job could just be triggered prematurely. This can easily be prevented by adding a dependency from the last computation job to the wake-up job. If the computation somehow flounders, the SAFETY_TIMEOUT (5s) will eventually raise an exception to let the test fail cleanly (shutting down the Scheduler automatically)	2023-12-13 22:55:28 +01:00
Ichthyostega	206c67cc8a	Scheduler-test: adapt λ-post to include deadline info ...it seems impossible to solve this conundrum other than by opening a path to override a contextual deadline setting from within the core Activity-Language logic. This will be used in two cases - when processing a explicitly coded POST (using deadline from the POST) - after successfully opening a Gate by NOTIFY (using deadline from Gate) All other cases can now supply Time::NEVER, thereby indicating that the processing layer shall use contextual information (intersection of the time intervals)	2023-12-13 19:42:38 +01:00
Ichthyostega	3bf3ca095b	Scheduler-test: failure of extended cascading notifications ...this is an interesting test failure, which highlights inconsistencies with handling of deadlines when processing follow-up from NOTIFY-triggers There was also some fuzziness related to the ''meaning'' of λ-post, leading to at least one superfluous POST invocation for each propagation; fixing this does not solve the problem yet removes unnecessary overhead and lock-contention	2023-12-13 19:27:45 +01:00
Ichthyostega	fcde92a476	Scheduler-test: add node-weight statistics ...playing around with the graph for the Scheduler integration test ...single threaded run time seemed to behave irregular ...but in fact it is very close to what can be expected based on an ''averaged node weight'' Fortunately its very simple to add that into the existing node statistics	2023-12-12 20:51:31 +01:00
Ichthyostega	eef3525710	Scheduler-test: setup for integration test Basically this is all done and settled already: this is the `usageExample()` from `TestChainLoadTest`. However, the focus is slightly different here: We want a demonstration that the Scheduler can work flawlessly through a massive load. Thus the plan is to use much more challenging parameters, and then lean back and watch what happens....	2023-12-12 19:21:15 +01:00
Ichthyostega	b987aa2446	Scheduler-test: single invocation of a computation load ...can now be assembled easily from existing parts ...use this setup as the simple introductory example in SchedulerService_test	2023-12-12 18:17:03 +01:00
Ichthyostega	23a3a274ce	Scheduler-test: investigate further breakage ...which turns out to be due to the DUMP-Statements, which seem to create quite some contention on their own. Test cases with very tight schedule will slip away then; without print statement everything is GREEN now	2023-12-12 01:55:52 +01:00
Ichthyostega	69faaef8cd	Scheduler-test: --- instrumentation --- This partially reverts commit `72f11549e6`. "Chain-Load: Scheduler instrumentation for observation" Hint: revert this changeset to re-introduce the print statements for diagnostic	2023-12-11 23:55:55 +01:00
Ichthyostega	da57e3dfcd	Scheduler-test: ''can demonstrate running a synthetic load'' (closes #1346 ) * added benchmark over synchronous execution as point of reference * verified running times and execution pattern * Scheduler behaves as expected for this example	2023-12-11 23:53:25 +01:00
Ichthyostega	347b9b24be	Scheduler-test: complete and integrate computational load This basically completes the Chain-Load framework; a simple Scheduler integration run with all relevant features can now be demonstrated.	2023-12-11 19:42:23 +01:00
Ichthyostega	db1ff7ded7	Scheduler-test: incremental calibration of both variants - Generally speaking, the calibration uses current baseline settings; - There are now two different load generation methods, thus both must be calibrated - Performance contains some socked and non-linear effects, thus calibration should be done close to the work point, which can be achieved by incremental calibration until the error is < 5% Interestingly, longer time-base values run slightly faster than predicted, which is consistent with the expectation (socket cost). And using a larger memory block increases time values, which is also plausible, since cache effects will be diminishing	2023-12-11 04:43:05 +01:00
Ichthyostega	9ef8e78459	Scheduler-test: implement memory-accessing load ...use an array of volatiles, and repeatedly add neighbouring cells ...bake the base allocation size configurable, and tie the alloc to the scale-step	2023-12-11 03:13:28 +01:00
Ichthyostega	df4ee5e9c1	Scheduler-test: implement pure computation load ..initial gauging is a tricky subject, since existing computer's performance spans a wide scale Allowing - pre calibration -98% .. +190% - single run ±20% - benchmark ±5%	2023-12-11 03:10:42 +01:00
Ichthyostega	beebf51ac7	Scheduler-test: draft a configurable CPU load component ...which can be deliberately attached (or not attached) to the individual node invocation functor, allowing to study the effect of actual load vs. zero-load and worker contention	2023-12-10 19:58:18 +01:00
Ichthyostega	fcfdf97853	Chain-Load: prepare infrastructure for computational load Within Chain-Load, the infrastructure to add this crucial feature is minimal: each node gets a `weight` parameter, which is assigned using another RandomDraw-Rule (by default `weight==0`). The actual computation load will be developed as a separate component and tied in from the node calculation job functor.	2023-12-09 03:13:48 +01:00

1 2 3 4 5 ...

2767 commits