aboutsummaryrefslogtreecommitdiff
path: root/lib/xixanta/src/assembler.rs
Commit message (Collapse)AuthorAgeFilesLines
* Ensure expressions in arguments are fully parsedMiquel Sabaté Solà2026-04-231-0/+13
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | Imagine the following scenario: lda #.lobyte(@address + 1) Here we have an immediate which comes from the .lobyte call. That being said, since this is using the '#' symbol, we look ahead after noticing it just in case we need to disambiguate some scenarios. When a parenthesis is found, in extract_paren_expression() we temporarily discard the looking ahead status as expressions inside of parenthesis are no longer affected by any hypothetical external ambiguity. This, though, was not being done in parse_arguments(), which is the function that would be called in the above scenario. Hence, ensure that we temporarily decrease the look ahead status for the inner arguments, and then increase it back when we are done. This way, when the parser finds '@address', it will no longer go the "we have found a 'value', refrain from ambiguities and return early" route, and it will check for any binary expression. Note that all this dance actually applied here, as binary operations are a source of ambiguities, and the look ahead mechanism was brought about because of these kinds of expressions. Hence, it's not like the look ahead mechanism is troublesome in on itself, we just missed the special handling on this specific scenario. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Implement the asan:stack statementMiquel Sabaté Solà2026-04-081-0/+116
| | | | | | | | | | | | | | This statement allows programmers to reserve a memory range for the stack. This is useful to reserve a memory region which is not attached to any variable, and: 1. It is a way to safely reserve space on the $0100 page. 2. The --stats/--write-info flags will be able to add up the memory space reserved for the stack. 3. Future tooling might be able to use this value as a way to detect stack overflows. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Re-create the values for bundle call argumentsMiquel Sabaté Solà2026-03-101-4/+38
| | | | | | | | | | | | | | | | | | Bundle call arguments were fine most of the times, when arguments could be processed as-is and there were no issues with arguments being overwritten by successive calls. This was not the case, though, whenever a given instruction was delayed into a PendingNode status. In this case, the argument would get the last value, and in some extreme cases that definition might not have been there any more. Prevent all of this by providing a list of PendingDefine's, which are a way to re-create these call-only arguments as they were initially found. These PendingDefine's are then created on "crunch" on each PendingNode. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* asan: ignore names from macro argumentsMiquel Sabaté Solà2026-03-091-4/+4
| | | | | | | These are usually capitalized, and thus they are not going to be following the usual zp_/m_/wr_ pattern. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Fix small indentation typoMiquel Sabaté Solà2026-02-251-1/+1
| | | | Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Allow using a byte for indirect jumpsMiquel Sabaté Solà2026-02-111-11/+12
| | | | | | | They should be expanded to a 16-bit address from the zero-page. The fact that the assembler was complaining about it was simply erroneous. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* asan: Fix validation on indirect jumpsMiquel Sabaté Solà2026-02-111-3/+1
| | | | | | | | | | | | | | There were two bugs involved at the same time. First, for some reason, the "matches" clause was negated, which defeats the purpose of the check. This even resulted in a bad test run which was accepted because I just did not caught it before. Second, pure indirect addressing mode is only available for the "jmp" instruction. Hence, there's no need for that "matches" at all, and even less to filter that based on "jsr" which doesn't even implement this addressing mode. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* asan: don't complain on arithmetic with addressesMiquel Sabaté Solà2026-02-061-4/+17
| | | | | | | This was already the case for plain addresses, but it was not being considered in the case of an arithmetic operation. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Fix the arithmetic on indexed addressing modesMiquel Sabaté Solà2026-02-051-51/+136
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | In instructions like: lda address + 1, x The assembler was picking the size from the "+ 1" side, making this part of a one-byte size and, hence, assuming it was zeropage indirect addressing instead of an absolute one. This meant that one byte was cut off the end binary, and the address was wrong altogether. Fix this by telling the assembler to assume a .size = 2; for variables which have been defined as addresses. Then, the operation evaluation can take that as an indication to set the size to 2 as well and, hence, assume absolute indirect addressing. Moreover, this bug was intertwined with another one, where arithmetic could be borked when both operands were of an unexpected size. For example: lda #(300 - 260) Here the size would've been resolved to 2 as well, but it would have also reported an error because both operands were 2-bytes long, despite the end result fitting on a single byte. Since this is an explicit immediate addressing, then it would have assumed the programmer to try that with a 16-bit value, which is wrong. This commit also fixes this bug as it was deeply intertwined to how numbers are computed on operations and how sizes are evaluated. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Make the fallthrough error more explicitMiquel Sabaté Solà2026-02-031-1/+4
| | | | | | | Add the target and effective addresses into the message, so it's more clear how different they are. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Avoid invalid identifiers in proc/macro/scopeMiquel Sabaté Solà2026-02-031-7/+58
| | | | | | | | This was apparently neglected and you were able to pick invalid identifiers to identify procs, macros and scopes. Ensure this does not happen again and provide tests for it. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Change .fallthrough to __fallthrough__Miquel Sabaté Solà2026-02-021-25/+18
| | | | | | | | | | | | | Implementing it as a control statement has the bad thing that it's impossible to be compatible with other assemblers such as ca65, as you cannot create dummy control statements or something like that in there. Instead of that, we define it with the special underscores which are still valid for identifiers, and give a "compiler-specific thingie" flair to it. This also has the benefit that the parser can be a bit more strict. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Add the .fallthrough control statementMiquel Sabaté Solà2026-02-021-1/+86
| | | | | | | | | | | | | | | | This is exclusive to 'nasm' and it allows the developer to explicitly tell the assembler than a "fall through" condition is actually desired: it's not a mistake. This comes in two flavors. The first, without arguments, just makes this explicit without much enforcement. The second allows you to pass an argument which is the name of the function or label you are expecting to fall through. The assembler will error out if the fall through address is not the expected one, hence telling the programmer whenever the fall through condition they thought in the past is no longer true (e.g. the function has moved somewhere else in the code). Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Warn on pointless (un)conditional branchingMiquel Sabaté Solà2026-02-021-0/+76
| | | | | | | | | | | | | | | | | | | | | | | | | | Sometimes performing some 'jmp'/'jsr' can be quite pointless, and the programmer might not be fully aware of this because of the layout of the code. Imagine: .proc foo ;; code jmp bar .endproc ;; Documentation, comments, extra space, etc. .proc bar ;; whatever .endproc The 'jmp' in the code above tries to perform a call stack optimization, but it's actually not needed because the next instruction after 'jmp' is the one inside of 'bar', but that's obfuscated because of the layout. In these sort of cases (and also for 'jsr' and branches) warn the programmer about it so it can remove that instruction. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* xixanta: add context from macro expansionsMiquel Sabaté Solà2026-02-021-1/+118
| | | | | | | | | | | | | The way macro expansions work is that the evaluated bundles replace the current node, but this throws away the context of the line, the source, etc. from the original line of code. Add a stack that is pushed/popped when expanding macros, and are then cloned for each pending node and error. This way, errors have full context of the original code and can display backtraces for macro expansion. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* asan: move the conflict check into its own loopMiquel Sabaté Solà2025-12-151-18/+30
| | | | | | | | | | | | | There was a condition race in which some conflicts would be caught but sometimes wouldn't depending on how/when `memory.memory_ranges` was being filled. Instead of this, just fill this vector with otherwise valid candidates (other checks like "is it used?" still apply in this context), and then perform the conflict check upon the already filled vector. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Allow constants in asan:reserve statementsMiquel Sabaté Solà2025-12-151-0/+16
| | | | | | | | | | | | | | | | This way, if you define a constant like: MY_BUFFER_LEN_IN_BYTES = $10 You can then declare your buffer like so: zp_buffer = $00 ; asan:reserve MY_BUFFER_LEN_IN_BYTES And then further in the code you can rely on just using the constant for bound checking, and then the address sanitizer will check on bound checks via static analysis as well. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* asan: addresses >= 0x2000 cannot be reserved spaceMiquel Sabaté Solà2025-12-151-10/+15
| | | | | | | These addresses are merely references to real APU/PPU addresses which can be used for ease of use. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Report on multi-byte values on arithmetic evaluationMiquel Sabaté Solà2025-12-151-0/+32
| | | | | | | | | | | | | | | When evaluating an arithmetic operation, check on whether it theoretically spans more than one byte, and set it on the resulting Bundle.size. It's up to the caller to decide on whether that makes sense or not. This fixes situations in which: sta $200 + 1 was marked as having 2 bytes for the size as only the "1" was being considered for the size on this operation. This would then signal the "sta" evaluation that it was zeropage instead of absolute addressing. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* Allow asan:reserve to receive an usizeMiquel Sabaté Solà2025-12-141-8/+8
| | | | | | | This allows for reserving memory regions which go outside of the page boundary. Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
* xixanta: Don't allow negative variables on asanMiquel Sabaté Solà2025-09-121-4/+41
| | | | | | | | | | | | | | Negative variables make no sense when it comes to do stuff like: variable = -1 lda variable That being said, as an assembler you never know the hacks and nonsense programmers are willing to endure. But we do know that if the address sanitizer is enabled, since if it's an "asan-friendly", then an explicit negative variable can be barred. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* asan: Add fixes on absolute/indirect addressingMiquel Sabaté Solà2025-09-031-1/+6
| | | | | | The check was not being applied on certain conditions. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Apply asan:ignore also when bundlingMiquel Sabaté Solà2025-09-031-1/+13
| | | | | | | This way instructions which might make use of bare memory numbers can freely ignore the address sanitizer when it actually makes sense. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Add a check for variable namesMiquel Sabaté Solà2025-09-031-0/+67
| | | | | | | This check ensures that asan-friendly names actually match their expected scope. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Validate that memory access is done via variablesMiquel Sabaté Solà2025-09-031-9/+127
| | | | | | | | | | | | | The address sanitizer is now able to detect whenever in an instruction a memory access is done without using variables. This is now detected for all instructions except for branching, which falls outside of this scope. Moreover, simple arithmetics is allowed and bounds are checked for simple cases. That being said, more involved bound checks should be done with other tools (e.g. emulators). Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Add a warning for each unused variableMiquel Sabaté Solà2025-09-021-21/+45
| | | | Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* assembler: Preserve the original working directoryMiquel Sabaté Solà2025-09-021-0/+10
| | | | Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* assembler: Add support for asan:reserve,ignoreMiquel Sabaté Solà2025-09-021-12/+213
| | | | | | | | This is the initial support for both directives for the address sanitizer. Note that asan:weak has been moved into asan:ignore, which is not exactly the same but for now it should suffice. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* nasm: Error out when using WRAM when not availableMiquel Sabaté Solà2025-08-281-0/+26
| | | | | | | | | | This commit introduces the ability to inspect the temptative header before producing the actual output, and with that it checks whether the Working RAM is being advertised or not. If it is not being advertised but the assembler detected memory accesses to that region, then we are in trouble and we should error out. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* xixanta: Be more mindful on what's being exportedMiquel Sabaté Solà2025-08-271-6/+21
| | | | | | | This means removing a lot of `pub` structs or enums, as well as adding documentation on `pub` structs. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* nasm: Implement the -s/--stats optionMiquel Sabaté Solà2025-08-271-0/+9
| | | | | | | | This option prints further information on how segments are laid out. In particular, for now it prints the amount of space being filled for each segment. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Don't get into all blocks when evaluating the contextMiquel Sabaté Solà2025-08-181-4/+51
| | | | | | | | For some control statements like .if/.ifdef/.ifndef this is only desired when the condition is true; otherwise getting into the inner block should be prevented. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Provide a node to the Object structMiquel Sabaté Solà2025-08-181-1/+59
| | | | | | | | | This allows the `evaluate_variable` to pull from it in the crunching stage so to evaluate the original node in cases like macro expansion, where the connection between the macro argument and the original caller might have been lost. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Reset literal mode before left arm of an operationMiquel Sabaté Solà2025-08-181-0/+5
| | | | | | | | On an arithmetical/logical operation, the literal mode needed to be reset before evaluating the left arm since the previous evaluation of the node could have altered it. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Fix shift overflow on certain expressionsMiquel Sabaté Solà2025-08-181-2/+2
| | | | | | | | | I got a report from libfuzz that some cryptic input could make the assembler panic on shifts. It turns out that the check on whether the operator was too big or not had to be explicitely casted to `usize` to avoid signedness issues. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Use a buffered reader for incbinMiquel Sabaté Solà2025-08-171-2/+3
| | | | | | | | | Iterating via `bytes()` on a file is inefficient as the default implementation calls `read` on each byte, which can be costly on bytes which are not in memory like files. This is extra important for statements like `incbin` as included files can be rather big. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Apply suggested style fixes from clippyMiquel Sabaté Solà2025-08-171-25/+17
| | | | Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Implement echo control statementsMiquel Sabaté Solà2025-01-221-1/+30
| | | | Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* nasm: Implement the -D flagMiquel Sabaté Solà2025-01-221-4/+64
| | | | | | | This allows users to define variables directly from the command line, which is useful for testing purposes. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Introduce looking ahead when parsing literalsMiquel Sabaté Solà2025-01-211-1/+1
| | | | | | | | | | | | | | | | | | | | | | There were certain situations in which syntax ambiguity could arise. For example, the parser as it stood could treat a valid expression such as 'lda #$80 >> 2' in an unexpected 'lda #$(80 >> 2)'. This is of course bad, and it came from the fact that literals don't have enclosing characters, and white spaces are not enough to provide disambiguation in some cases. Because of this fact, the parser now has a "look ahead" capability similar to many other parsers, and it's applied for now only to literals. This looking ahead actually honors parenthesis, so these can be added if the programmer wants to explicitely disambiguate an expression. This involved quite the heavy lifting, and some functions like 'parse_expression_with_identifier' had to be removed with the rewrite. This had the side effect of having (hopefully) more sane functions all around, and the parser also has a better capability to differentiate between regular Values and Calls. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Whitelist valid identifier charactersMiquel Sabaté Solà2025-01-161-5/+14
| | | | | | | | Instead of making up a list of characters that end an identifier, do the other way around since it's far less cumbersome and it prevents from silly bugs such as "var+1" being considered a single identifier. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Permit 16-bit decimal valuesMiquel Sabaté Solà2025-01-161-9/+18
| | | | | | | | | | As a remnant of old code, the 'parse_decimal' function was not allowing for decimal values larger than 8-bits. This was not the case in other areas such as 'parse_hexadecimal', and in the rest of the code we already cover that immediates are not too big in instructions. Hence, this restriction can be lift up and allow up to 16-bit decimal literals. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Implement .ifdef/.ifndef statementsMiquel Sabaté Solà2025-01-161-0/+49
| | | | | | | They are just synonyms for ".if .defined" and ".if !.defined" respectively. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Implement .def/.defined expressionsMiquel Sabaté Solà2025-01-161-0/+53
| | | | | | | This can be combined with .if/.elsif statements just like any other expression. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Implement .if/.elsif/.else statementsMiquel Sabaté Solà2025-01-161-11/+133
| | | | Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Add boolean operatorsMiquel Sabaté Solà2025-01-161-7/+60
| | | | | | | | | | | Allow for expressions that evaluate to a boolean expression. This in turn mean that the value is just set to 0 or 1 depending on the given condition. As with other assemblers, only a value of 0 evaluates to 0, and others go to 1. So, something like "1 && 2" evaluates to 1 even if it doesn't make much sense at first glance (as an assembler we just assume that the programmer knows what it's doing). Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Move check for ASCII-only strings into the parserMiquel Sabaté Solà2025-01-151-9/+2
| | | | Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Implement .asciiz control statementMiquel Sabaté Solà2025-01-141-10/+69
| | | | Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* Implement the .res control statementMiquel Sabaté Solà2025-01-141-0/+97
| | | | | | | | | | | | | Similarly to other assemblers, this allows the programmer to write a definite amount of bytes with the same values. Compared to other assemblers there are two things to notice. First, there is a limit to it (i.e. whatever can fit in 2 bytes). Second, if the fill value is not provided, then it will default to the current mapping's fill value, or just 0x00 if the current mapping doesn't define one of its own. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
* tests: Update code.nes so -Werror can be passedMiquel Sabaté Solà2025-01-131-20/+2
| | | | | | | | | Update the underlying code.nes test data so it uses a tighter configuration for NROM which in turn frees nasm from complaining about empty segments. This makes room to allo for -Werror so we catch warnings that might appear in the future. Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>